Faq Shortcuts
kishormorol/cli-faq-shortcuts
Turn the questions and instructions a user keeps typing into a project into short slash-command skills, mined from their real Claude Code, Codex and Cursor session history and git log.
Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage &…
$ npx skills add microsoft/eval-guide --skill eval-faq -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/eval-guide eval-faq --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-faq .claude/skills/eval-faq && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-faq" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-faq into .claude/skills/eval-faq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-faq", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/eval-guide/tree/main/skills/eval-faqType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/eval-guide --skill eval-faq -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/eval-guide eval-faq --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-faq .agents/skills/eval-faq && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-faq" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-faq into .agents/skills/eval-faq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-faq", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/eval-guide --skill eval-faq -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/eval-guide eval-faq --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-faq .cursor/skills/eval-faq && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-faq" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-faq into .cursor/skills/eval-faq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-faq", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/eval-guide.git --path skills/eval-faq--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/eval-guide --skill eval-faq -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/eval-guide eval-faq --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-faq .gemini/skills/eval-faq && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-faq" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-faq into .gemini/skills/eval-faq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-faq", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/eval-guide eval-faqInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/eval-guide --skill eval-faq -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-faq .github/skills/eval-faq && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-faq" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-faq into .github/skills/eval-faq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-faq", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/eval-guide --skill eval-faq -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/eval-guide eval-faq --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-faq .opencode/skills/eval-faq && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-faq" agent skill from https://github.com/microsoft/eval-guide/tree/main/skills/eval-faq into .opencode/skills/eval-faq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-faq", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-faqAnswers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage &…
Eval Faq is an agent skill from microsoft/eval-guide, published by the product's own GitHub organization. Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage & Improvement Playbook, Eval Guidance Kit) supplemented by select industry sources.
Its SKILL.md is about 10k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Sales & Support, covering Agent evaluation and testing and Help center and FAQ content. It works with Microsoft Copilot Studio. The repository describes itself as: A plugin for AI agent evaluation. Plan evals, generate test cases, interpret results for Copilot Studio agents. Grounded in Microsoft's Eval Scenario Library & Triage Playbook. The licence is MIT.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 7a22a89. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
learn.microsoft.comgithub.comanthropic.comaka.mseugeneyan.comhamel.devbraintrust.devFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Faq loads about 10k tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 4,805 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from microsoft/eval-guide at commit 7a22a89, republished under its MIT licence (© microsoft). 4,805 words, ~10,464 tokens.
.claude/skills/eval-faq/SKILL.md (or your agent's skills folder).Answer any question about eval methodology, grader types, dataset design, criteria writing, non-determinism, tool-call evaluation, multi-turn agent evaluation, eval tooling, capability vs. regression evals, and interpreting results — specifically in the context of AI agent evaluation. The primary methodology is skills/eval-guide/playbook.md: Practical Guidance on Agent Evaluation: a 10-step playbook. Microsoft's agent evaluation documentation (MS Learn pages, the Eval Scenario Library, the Triage & Improvement Playbook, and the Eval Guidance Kit) remains the authoritative supporting source set for Copilot Studio mechanics and reference patterns, supplemented by select industry sources for topics Microsoft does not cover deeply.
When invoked as /eval-faq <question>, follow this process exactly:
Use this topic-to-URL routing table to decide what to fetch. Fetch FIRST, then answer. Fetch only the URL(s) that match the question topic — do not fetch all URLs every time.
| Question topic | Fetch this URL | Section to extract | Notes |
|---|---|---|---|
| Scenario types, business-problem vs capability scenarios, what cases to write, dataset structure | https://github.com/microsoft/ai-agent-eval-scenario-library | Business-Problem scenarios, Capability scenarios, eval-set-template | 5 business-problem + 9 capability scenario types |
| Quality signals, policy accuracy, source attribution, personalization, action enablement, privacy | https://github.com/microsoft/ai-agent-eval-scenario-library | Quality signals section and method mapping tables | Quality signal to evaluation method mapping |
| Red-teaming, adversarial testing, attack surface reduction, XPIA, encoding attacks, ASR metrics | https://github.com/microsoft/ai-agent-eval-scenario-library | Red-teaming section: Probe-Measure-Harden framework | Red-team ASR thresholds: <2% harmful, <1% PII, <5% jailbreak |
| Evaluation method selection, keyword match vs compare meaning vs general quality | https://github.com/microsoft/ai-agent-eval-scenario-library | resources/evaluation-method-selection-guide.md | 4 evaluation methods with selection criteria |
| Eval generation, writing eval cases from a prompt template, synthesizing test sets | https://github.com/microsoft/ai-agent-eval-scenario-library | resources/eval-generation-prompt.md | Template for generating eval cases |
| Agent profile template, defining agent scope for eval | https://github.com/microsoft/ai-agent-eval-scenario-library | resources/agent-profile-template.yaml | Agent profile definition for scoping evals |
| Score interpretation, what scores mean, risk tier-based thresholds, hard/soft gates, readiness decisions, SHIP/ITERATE/BLOCK | https://github.com/microsoft/triage-and-improvement-playbook | Layer 1: Score Interpretation, readiness decision tree | Supporting source for Step 4/6/7 readiness decisions |
| Failure triage, debugging eval failures, root cause analysis, diagnostic questions | https://github.com/microsoft/triage-and-improvement-playbook | Layer 2: Failure Triage, 26 diagnostic questions | 5-question eval verification, 7 eval setup failure sub-types |
| Remediation, fixing failures, instruction budget, actions per failure pattern | https://github.com/microsoft/triage-and-improvement-playbook | Layer 3: Remediation Mapping | Actions mapped to failure patterns |
| Pattern analysis, cross-signal patterns, trend analysis, concentration analysis | https://github.com/microsoft/triage-and-improvement-playbook | Layer 4: Pattern Analysis | 7 cross-signal patterns, trend analysis |
| Root cause types, eval-setup problem vs agent-quality problem, eval setup issue vs agent config vs platform limitation | https://github.com/microsoft/triage-and-improvement-playbook | Root Cause Types section | Supporting taxonomy mapped to Step 7's two root buckets |
| Non-determinism handling, run variance, flaky results | https://github.com/microsoft/triage-and-improvement-playbook | Non-determinism section | 3 runs minimum, +/-5% normal, +/-10% investigate |
| 4-stage iterative framework, Define, Set Baseline & Iterate, Systematic Expansion, Operationalize | https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/evaluation-iterative-framework | Full framework — all 4 stages | Supporting MS Learn lifecycle/cadence source under the 10-step playbook |
| Eval checklist, readiness checklist, pre-launch verification | https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/evaluation-checklist | Full checklist | Maps to Eval Guidance Kit documents |
| Grader types, code-based vs LLM-judge vs human graders, common evaluation approaches | https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/architecture/common-evaluation-approaches | Echo, Historical Replay, Synthesized Personas; grader types | 3 approaches + 3 grader categories |
| 7 test methods, General Quality, Compare Meaning, Capability Use, Keyword Match, Text Similarity, Exact Match, Custom | https://learn.microsoft.com/en-us/microsoft-copilot-studio/analytics-agent-evaluation-overview | 7 test methods section | General Quality sub-dimensions: Relevance, Groundedness, Completeness, Abstention |
| Test set creation, building eval datasets in Copilot Studio | https://learn.microsoft.com/en-us/microsoft-copilot-studio/analytics-agent-evaluation-create | Test set creation methods | Generate, import, or manually write test cases |
| Test set editing, user profiles, connections, modifying test methods | https://learn.microsoft.com/en-us/microsoft-copilot-studio/analytics-agent-evaluation-edit | Manage user profiles and connections, edit test methods | Multi-profile eval for simulating different users; GCC limitations |
| Running evals, viewing results, test results interpretation | https://learn.microsoft.com/en-us/microsoft-copilot-studio/analytics-agent-evaluation-results | Run tests and view results | 89-day result retention; export results immediately |
| Agent evaluation overview, why use automated testing, test chat vs eval | https://learn.microsoft.com/en-us/microsoft-copilot-studio/analytics-agent-evaluation-intro | About agent evaluation | GCC limitations: no user profiles, no Text similarity method |
| Rubric refinement workflow, aligning AI grading with human judgment | https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/kit-rubrics-refinement-workflow | 8-step workflow: Run, Review, Grade, Refine, Save, Re-run, Repeat | Alignment matrix, Standard vs Full refinement views, example marking |
| Rubric best practices, tips for rubric refinement | https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/kit-rubrics-best-practices | Best practices for refinement | Quality over quantity for examples; don't chase 100% alignment |
| Rubric reference guide, grade definitions, rubric structure | https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/kit-rubrics-reference | Rubrics reference | Grade scale definitions, rubric components |
| Copilot Studio Kit overview, kit capabilities | https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/kit-overview | Kit overview | Parent page for all Kit features including rubrics |
| 11 scenario validation themes, evaluation frameworks | https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/architecture/evaluation-frameworks | 11 scenario validation themes | |
| Defining eval purpose, what to evaluate, scoping eval | https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/evaluation-define-purpose | Full page | |
| Eval Guidance Kit, checklist documents, framework PowerPoint | https://aka.ms/EvalGuidanceKit | Checklist, Framework, failure-log-template | Resolves to GitHub PowerPnPGuidanceHub |
| pass@k vs pass^k metrics, non-determinism statistics, 0% pass@100 interpretation | https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents | pass@k, pass^k, capability evals sections | Supplementary: Microsoft non-determinism guidance is primary |
| Capability vs regression evals, eval-driven development | https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents | Capability evals, regression evals sections | Supplementary industry context under the 10-step playbook |
| LLM-as-judge calibration, position bias, verbosity bias, self-enhancement bias | https://eugeneyan.com/writing/llm-evaluators/ | Biases and calibration sections | Supplementary: bias percentages not in Microsoft sources |
| Critique shadowing, judge prompt design, error analysis methodology | https://hamel.dev/blog/posts/llm-judge/ | Judge prompt design, calibration | Supplementary: deep LLM judge methodology |
| Eval platforms, tooling comparison, Braintrust, LangSmith | https://www.braintrust.dev/articles/top-5-platforms-agent-evals-2025 | Platform comparison | Supplementary: lightweight tooling reference |
| Any question not clearly matching above | Fetch https://learn.microsoft.com/en-us/microsoft-copilot-studio/guidance/evaluation-overview as primary source, supplement with relevant knowledge base section | Default fallback is MS Learn |
Fetch rules:
Synthesize the fetched content with the knowledge base below. The 10-step playbook is the methodology spine; Microsoft fetched content supplies supporting details and Copilot Studio specifics, then external sources fill gaps.
Answer style rules — no exceptions:
Use the sections below as your primary reference when fetched content does not cover the question, or to supplement fetched content with additional details.
The core methodology is skills/eval-guide/playbook.md: Practical Guidance on Agent Evaluation: a 10-step playbook. Use the MS Learn pages below as supporting sources, not the spine.
Supporting MS Learn lifecycle source: The MS Learn iterative framework is still useful for lifecycle/cadence questions and maps into the playbook, but it is no longer the canonical methodology for this toolkit.
Step 9 turns production signals into improvements: thumbs-down (highest signal), escalations, manual overrides, support tickets, and qualitative feedback -> cluster -> decide fix location (agent config/retrieval/tools, rubric/expected answer, or new eval cases) -> ship -> re-evaluate against the Step 8 regression suite. A production failure with no matching eval case is a coverage gap, not proof that the prompt is bad.
Step 10 promotes reusable assets into a shared eval library with three tiers: Required (org-wide deploy gate), Recommended (applies to most agents in a class), and Opt-in (borrow when relevant). Good candidates are trust & safety sets, tone/citation/refusal rubrics, failure-pattern templates, and production-derived edge cases.
Per Microsoft's Eval Scenario Library, scenarios divide into two categories:
5 Business-Problem scenarios (test whether the agent solves the real user problem):
9 Capability scenarios (test a specific isolated ability):
Anti-pattern: Skewing your dataset 80%+ toward happy-path cases. Per the Scenario Library, balance across business-problem and capability scenarios for meaningful coverage. Target roughly 50% happy-path, 30% edge cases, 20% adversarial.
Microsoft's Eval Scenario Library includes five reusable quality dimensions that can inform eval-set design, but the toolkit records them as workbook registry rows rather than treating them as the primary planning artifact:
Each dimension can map to methods such as Keyword Match, Compare Meaning, Capability Use, or General Quality, but the selected method and governance belongs to the eval set in the workbook registry.
Per MS Learn agent evaluation guidance, seven test methods cover different evaluation needs:
Per MS Learn (common-evaluation-approaches), three approaches for generating test interactions:
Per the Triage Playbook, score interpretation follows a 4-layer framework:
Layer 1 — Score Interpretation: Apply risk tier, workbook-defined gates, grader-validation caveats, and the readiness decision tree:
Layer 2 — Failure Triage: When scores are low, run the 5-question eval verification first (is the eval itself correct?) before blaming the agent. Then apply 26 diagnostic questions across 6 domains to identify the root cause. Seven eval setup failure sub-types cover common grader/dataset bugs.
Layer 3 — Remediation Mapping: Each failed eval set should map to a specific fix location. Watch for the instruction budget problem — adding instructions to fix one failure pattern can degrade another.
Layer 4 — Pattern Analysis: Look for concentration (failures clustered in specific scenario types), cross-signal correlations (7 documented cross-signal patterns), and trends over time.
Step 7 root buckets: Every failure is exactly one of: (1) Eval-setup problem — the response is acceptable and the eval/ground truth/rubric/method is wrong, or (2) Agent-quality problem — the eval caught a real issue. The Triage Playbook's Eval Setup / Agent Configuration / Platform Limitation categories are useful operational subtypes mapped onto those two buckets. Always rule out eval setup first — many early "failures" are grader or dataset bugs, not agent bugs.
Per the Triage Playbook: agents are non-deterministic. Run a minimum of 3 trials per case. Score variance of +/-5% across runs is normal. Variance of +/-10% or more requires investigation — either the eval is flaky or the agent has a genuine instability.
Additional industry context from Anthropic: pass@k ("succeeded at least once in k runs") vs. pass^k ("succeeded every time in k runs") diverge massively at scale. At k=10 with 70% per-trial success: pass@k is approximately 97%, pass^k is approximately 3%. The same agent looks excellent or catastrophic depending on which metric you report. For customer-facing agents, pass^k is the right question. A 0% pass@100 is almost always a task specification problem, not an agent problem — fix the task definition before blaming the model.
Per Microsoft's Eval Scenario Library, red-teaming uses the Probe-Measure-Harden framework:
Red-team thresholds: ASR <2% for harmful content, <1% for PII leakage, <5% for jailbreak. Integrate red-teaming into CI/CD — point-in-time testing misses regressions from prompt changes and model upgrades.
Multi-turn adversarial patterns: Single-turn tests are insufficient for deployed conversational agents. Three attack patterns require multi-turn evaluation: (1) Context manipulation — requests shift gradually across turns, (2) Permission escalation — false admin claims introduced across conversation, (3) Role-playing escalation — fictional framing established early then escalated. Include at least 2-3 multi-turn adversarial scenarios in any eval suite.
Per MS Learn (common-evaluation-approaches), three grader categories:
Grading hierarchy (cheapest to most expensive): Run code-based checks first, then LLM judges on passing cases, then human review on a calibration sample. Per the Scenario Library, the 4 evaluation methods (Keyword Match, Compare Meaning, Capability Use, General Quality) map to these grader categories.
Calibration threshold: If your LLM judge and a human expert agree on fewer than 80% of cases (kappa < 0.6), your criteria are ambiguous. Rewrite criteria before trusting scores.
Per the Eval Scenario Library, use the eval-set-template.md to structure your dataset. Use the eval-generation-prompt.md template to generate cases from an agent profile.
agent-profile-template.yaml) to define scope before writing cases.CSV and scoring conventions: Copilot Studio import CSVs are exactly two columns: Question, Expected response. Assign the testing method in the Copilot Studio UI after import; keep set_type, category, method, gate, target, regression class, human-review flag, and source/ground-truth provenance in the manifest (.docx report + stage-N-data.json). Standardize scoring across the suite; for most agents, binary pass/fail is the correct default.
Per the 10-step playbook, evaluation starts at Step 1 — Plan the eval effort before the agent is built:
Anti-pattern: Writing evals after building the feature. That produces evals calibrated to what you built, not what you intended.
Per Step 7 and the Triage Playbook (Layer 2), never trust a score you have not manually verified. The first question is whether the failure is an eval-setup problem: Is the test set correct? Is the grader measuring the right thing? Is the expected answer actually right? Is the agent getting the right context? Is the eval environment matching production?
Axial coding process for failure analysis:
Per Step 7, always include "eval-setup problem" as a category — many failures in a new eval are grader, rubric, stale ground-truth, or manifest bugs rather than agent-quality problems.
Additional industry context from Hamel Husain: The axial coding methodology and "highest ROI activity in AI engineering" framing come from Hamel Husain's error analysis work. His key insight: most practitioners skip categorization and jump to "fix the prompt," missing structural patterns.
Per the Eval Scenario Library's Tool Invocations capability scenario and MS Learn's Capability Use test method:
Per MS Learn's evaluation approaches, multi-turn workflows require conversation-level evaluation, not turn-level:
Per MS Learn's evaluation frameworks (11 scenario validation themes):
The evaluation approach differs significantly based on agent complexity:
| Dimension | Simple Q&A agent | Multi-step / agentic workflow |
|---|---|---|
| Primary metric | Response accuracy (Compare Meaning, General Quality) | Task completion — did the end-to-end job get done? |
| Grading unit | Single turn: one input, one output | Conversation or trajectory: full sequence of steps |
| Key eval-set focus | Grounded answers and policy accuracy | Action enablement, tool invocation, and Q&A correctness |
| Test method mix | Heavy on Compare Meaning + General Quality | Add Capability Use for tool calls, Keyword Match for intermediate checkpoints |
| Failure modes to watch | Wrong answer, hallucination, refusal | Compounding errors, wrong tool selection, unnecessary steps, partial completion |
| Edge cases | Ambiguous queries, out-of-scope questions | Mid-workflow failures, tool timeouts, user corrections mid-conversation |
| Eval complexity | Low — deterministic input/output pairs work well | High — must evaluate intermediate steps AND final outcome |
Practical guidance:
No single eval method catches every failure. Per the Eval Scenario Library's 4 evaluation methods and the Triage Playbook's multi-layer approach:
Per MS Learn's General Quality test method, LLM judges evaluate across sub-dimensions (Relevance, Groundedness, Completeness, Abstention). Calibrate judges against these defined dimensions.
Additional industry context from Eugene Yan (bias data):
Additional industry context from Hamel Husain (critique shadowing): When building LLM judges from scratch, use the 7-step Critique Shadowing methodology: (1) Identify one expert, (2) Create diverse dataset, (3) Collect binary pass/fail with written critiques, (4) Fix obvious errors, (5) Build judge prompts iteratively using expert examples, (6) Error analysis on disagreements, (7) Build specialized judges for specific failure modes. Target >90% agreement with domain expert before production use.
Per the Eval Scenario Library's Knowledge Grounding guidance:
Per Step 9 of the 10-step playbook, eval is not a pre-launch gate — it is a continuous optimization loop:
When the agent passes evals but fails in production: Per the Triage Playbook, this is almost always a distribution mismatch. Pull 20 recent production failures. Check whether any would fail against your current eval dataset. If none would, your dataset needs production cases, not a better prompt.
Per the Triage Playbook's readiness decision tree:
Per the Triage Playbook's Layer 4 (Pattern Analysis): look for failure concentration in specific scenario types, cross-signal correlations, and trends over time. When a grader's verdict disagrees with your intuition, investigate — either the grader is wrong (fix the criterion) or your intuition is wrong (update your mental model).
For tooling questions, the primary recommendation is Microsoft's Copilot Studio evaluation features for production Copilot agents. For teams needing third-party platforms:
After answering the question, check whether the user would benefit from running a sibling eval skill. If so, append a one-line recommendation at the end of your answer.
| If the question involves... | Suggest this skill | One-liner to append |
|---|---|---|
| Creating an eval plan or scoping what to evaluate | /eval-suite-planner | "For a populated Eval Suite Template workbook, run /eval-suite-planner." |
| Generating test cases, writing CSV datasets, building eval sets | /eval-generator | "To generate ready-to-import test case CSVs, run /eval-generator." |
| Interpreting scores, reading results, understanding pass rates | /eval-result-interpreter | "To interpret a specific set of eval results, paste them into /eval-result-interpreter." |
| Debugging failures, triaging low scores, root cause analysis, remediation | /eval-triage-and-improvement | "To triage specific failures with the full diagnostic framework, run /eval-triage-and-improvement." |
| What is eval, why eval matters, explaining eval to stakeholders | /eval-guide | "For an end-to-end eval explainer you can share with stakeholders, run /eval-guide." |
Rules:
/eval-faq (that is this skill — they are already here)./eval-faq What eval scenarios should I use for a RAG agent?
/eval-faq How do I interpret a 75% knowledge grounding score?
/eval-faq What is the difference between business-problem and capability scenarios?
/eval-faq When should I use a model-graded grader instead of a deterministic one?
/eval-faq What makes a good adversarial test case?
/eval-faq How many cases do I need in a dataset to get meaningful signal?
/eval-faq My eval passes 100% on first run — is that good?
/eval-faq How do I write a good criterion for a model-graded grader?
/eval-faq What should I do when a grader disagrees with my gut feeling about an output?
/eval-faq How do I handle non-determinism in my eval results?
/eval-faq My agent makes tool calls — how do I eval those?
/eval-faq I suspect my grader is wrong — how do I debug it?
/eval-faq What should I eval in production after I ship?
/eval-faq Should I use pass@k or pass^k for my agent?
/eval-faq How do I calibrate my LLM-as-judge grader?
/eval-faq When do I stop adding eval cases and just ship?
/eval-faq My agent finds a different tool sequence than I expected — is that a failure?
/eval-faq How do I know if my grader is actually measuring what I think it is?
/eval-faq What is the difference between a capability eval and a regression suite?
/eval-faq How do I eval a multi-turn conversational agent?
/eval-faq What eval platform or tool should I use?
/eval-faq My agent passes evals but fails in production — why?
/eval-faq How do I score intermediate steps in a multi-step agent?
/eval-faq How is evaluating a multi-step workflow different from a simple Q&A agent?
/eval-faq What does 0% pass@100 mean — is my agent broken?
/eval-faq How do I avoid LLM judge bias in my grader?
/eval-faq Which eval sets should I include?
/eval-faq What is the Probe-Measure-Harden red-teaming framework?
/eval-faq What are the 7 test methods in Copilot Studio?
/eval-faq How do I use the Triage Playbook to debug failing scores?
/eval-faq How does the MS Learn iterative framework relate to the 10-step playbook?
/eval-faq What are the 3 root cause types for eval failures?
/eval-faq How do I decide between SHIP, ITERATE, and BLOCK?
/eval-faq What red-team ASR thresholds should I target?
/eval-faq How do I generate eval cases from a prompt template?
/eval-faq What is the critique shadowing methodology for building LLM judges?
/eval-faq Should I use a 1-5 scale or pass/fail for my LLM judge?
/eval-faq How do I continuously red-team my agent in CI/CD?
/eval-faq How do I systematically analyze eval failures to find patterns?
/eval-faq How do I know if my eval is too easy?
/eval-faq How do I write an LLM grader prompt that actually works?
/eval-faq Should I score factuality and tone in the same eval criterion?
/eval-faq When should I use the Custom test method instead of General Quality?
/eval-faq How do I set up a Custom test method for compliance checking?© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/eval-faq of microsoft/eval-guide.
Open the folder on GitHubat commit 7a22a89
Eval Faq next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Faq this skillmicrosoft/eval-guide | 138 | — | ~10k | Automated safety check: Pass | MIT | |
| Faq Shortcutskishormorol/cli-faq-shortcuts | 127 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| OpenheardHeilonng23/openheard | 173 | — | ~1.5k | Automated safety check: Pass | AGPL-3.0 | |
| Write Docsbeyonders-studio/initiative | 170 | — | ~3.8k | Automated safety check: Pass | AGPL-3.0 | |
| Cc10x Guideromiluz13/cc10x | 164 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Docouture Writing Docs PagesInditexTech/weavejs | 208 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 |
kishormorol/cli-faq-shortcuts
Turn the questions and instructions a user keeps typing into a project into short slash-command skills, mined from their real Claude Code, Codex and Cursor session history and git log.
Heilonng23/openheard
Run an openheard feedback board through its MCP server - connect it, set up a workspace from a website, install the widget in this codebase, triage and dedupe feedback, reply, ship and announce…
beyonders-studio/initiative
Write or edit the Initiative help center under docs/en/ — the house voice, the structural rules, and how to build and check the site.
romiluz13/cc10x
Answers questions about cc10x itself — what it is, how to install and configure it, how the router, workflows, memory, and hooks operate, and how to troubleshoot.
InditexTech/weavejs
How to author AsciiDoc content in a docouture Antora documentation site — the content tree, xref: references, nav.adoc, admonitions, code blocks, and this site's own custom blocks (tabs, cards…
yennanliu/CS_basics
File an interview question and answer into doc/faq/ in the shape the other 49 FAQs use — the sheet its Scope line owns, the numbered section it belongs under, tagged code fences — and write the…
microsoft/eval-guide
A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps…
microsoft/eval-guide
Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.
microsoft/eval-guide
Generate standalone — turns the populated Eval Suite Planning workbook (output of /eval-suite-planner) into concrete capability eval sets and trust & safety eval sets.
microsoft/eval-guide
Analyzes Copilot Studio evaluation results using Practical Guidance on Agent Evaluation's 10-step playbook (Steps 6, 7, and 9) plus Microsoft's triage diagnostics.
microsoft/eval-guide
Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.
Works with
Categories
Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage &…. Eval Faq is an agent skill from microsoft/eval-guide, published by the product's own GitHub organization. Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage & Improvement Playbook, Eval Guidance Kit) supplemented by select industry sources.
Eval Faq fits situations like: tasks that involve Agent evaluation and testing; tasks that involve Help center and FAQ content.
Run `npx skills add microsoft/eval-guide --skill eval-faq -a claude-code`. Or copy the skill folder (skills/eval-faq in microsoft/eval-guide) into .claude/skills/eval-faq in your project. Claude Code loads it when a task matches its description.
Run `npx skills add microsoft/eval-guide --skill eval-faq -a codex`. Or copy the skill folder (skills/eval-faq in microsoft/eval-guide) into .agents/skills/eval-faq in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/eval-guide --skill eval-faq -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-faq, .gemini/skills/eval-faq, .github/skills/eval-faq and .opencode/skills/eval-faq in your project.
SKILL.md names no scripts, command-line tools or credentials: Eval Faq is instructions for the agent only.
SKILL.md names 7 domains. In commands or code: learn.microsoft.com, github.com, anthropic.com, aka.ms, eugeneyan.com, hamel.dev and braintrust.dev; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Faq is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 10k tokens (SKILL.md is roughly 42k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Faq: Faq Shortcuts (kishormorol/cli-faq-shortcuts, 127 stars), Openheard (Heilonng23/openheard, 173 stars), Write Docs (beyonders-studio/initiative, 170 stars) and Cc10x Guide (romiluz13/cc10x, 164 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/eval-guide, which has 138 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on June 24, 2026.
Source: microsoft/eval-guide on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.