Official agent skill

Eval Triage And Improvement

by microsoft in microsoft/eval-guide

A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps…

OfficialMITAuto-check passedTesting & QA

Install Eval Triage And Improvement

skills CLI
$ npx skills add microsoft/eval-guide --skill eval-triage-and-improvement -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/eval-guide eval-triage-and-improvement --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/eval-guide.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-triage-and-improvement .claude/skills/eval-triage-and-improvement && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-triage-and-improvement
GitHub stars
138
Token cost
~5.9k tokens
SKILL.md length
2,544 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps…

  • Works in 8 steps: Gather Eval Results → Score Interpretation → Pre-Triage Infrastructure Check → …
  • The users Copilot Studio agent evaluations have come back and they need to interpret scores
  • SKILL.md covers Workflow, Non-Determinism Handling, Step 9 Optimization Loop:… and Human Review Checkpoints, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Triage And Improvement is an agent skill from microsoft/eval-guide, published by the product's own GitHub organization. Use this skill when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps, or analyze patterns to improve their agent. Always use this skill when the user mentions: "eval failed", "why did this fail", "triage", "diagnose failure", "low pass rate", "fix evaluation results", "not passing", "failing test cases", "evaluation results", "improve my eval scores", or any situation where eval scores need…

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Agent evaluation and testing, Root cause analysis and Test generation. It works with Microsoft Copilot Studio and Microsoft Word. The repository describes itself as: A plugin for AI agent evaluation. Plan evals, generate test cases, interpret results for Copilot Studio agents. Grounded in Microsoft's Eval Scenario Library & Triage Playbook. The licence is MIT.

When your agent uses it

  • The users Copilot Studio agent evaluations have come back and they need to interpret scores
  • Diagnose root causes of underperforming test cases
  • Find remediation steps
  • Analyze patterns to improve their agent

Example prompts

  • “eval failed”
  • “why did this fail”
  • “triage”
  • “/eval-triage-and-improvement”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Gather Eval Results
  2. Score Interpretation
  3. Pre-Triage Infrastructure Check
  4. Prioritize Failures
  5. Classify Root Cause
  6. Map to Remediation
  7. Triage Rationale (teach the WHY)
  8. Generate Triage Report

What it can do on your machine

Read from SKILL.md and the folder at commit 7a22a89. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Triage And Improvement loads about 5.9k tokens when it runs. Until then it costs about 142 tokens; SKILL.md has 2,544 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~142
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/eval-guide at commit 7a22a89, republished under its MIT licence (© microsoft). 2,544 words, ~5,923 tokens.

Download SKILL.mdSave it as .claude/skills/eval-triage-and-improvement/SKILL.md (or your agent's skills folder).
name
eval-triage-and-improvement
description
Use this skill when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps, or analyze patterns to improve their agent. Always use this skill when the user mentions: "eval failed", "why did this fail", "triage", "diagnose failure", "low pass rate", "fix evaluation results", "not passing", "failing test cases", "evaluation results", "improve my eval scores", or any situation where eval scores need interpretation and action.

Eval Triage & Improvement

You help users interpret their agent evaluation results and find actionable next steps to improve. Follow the hybrid workflow: gather eval results first, then generate a structured triage report with Step 7 root buckets, owners, and recommended fixes.

This skill is grounded in skills/eval-guide/playbook.md, the canonical Practical Guidance on Agent Evaluation: 10-step playbook. It is the deep-dive for Step 7 — Iterate to Diagnose Failures and seeds Step 9 — Optimization Loop for production feedback. MS Learn pages and the Eval Guidance Kit remain supporting sources for Copilot Studio mechanics, lifecycle cadence, and checklist artifacts.

When to use this skill vs. eval-result-interpreter

These two skills share the same triage framework but serve different modes of work:

Use eval-triage-and-improvement when…Use eval-result-interpreter when…
You want interactive guidance walking through diagnosis step by stepYou have a CSV file or concrete results and want a one-shot structured report
You are in an ongoing improvement loop — fixing, re-running, and re-triagingThis is your first look at results — you need a verdict and top actions fast
You need detailed remediation help for specific eval-set failure patterns (e.g., "wrong tool fires — now what?")You want a customer-deliverable artifact (the .docx triage report)
You have many failures (15+) and need help prioritizing which to investigateThe eval run is relatively straightforward (<20 failures)
You need the playbook worked examples and deeper diagnostic walkthroughsYou need the activity map / result comparison tool recommendations inline

If in doubt: Start with eval-result-interpreter to get the structured report, then switch to eval-triage-and-improvement if you need interactive help implementing the fixes.

Workflow

Step 1: Gather Eval Results

Ask the user to share:

  1. Which eval sets ran and their pass rates (e.g., "Faithfulness: 71%, Prompt injection: 95%")
  2. Methodology manifest metadata from the companion .docx report or stage-N-data.json: set_type, category/capability dimension, method, gate (hard/soft), target, regression_class, human-review flag, and source/ground-truth provenance
  3. Specific failing test cases — the test case ID, sample input, expected value, actual agent response, and eval method assigned in Copilot Studio
  4. How many times they've run — is this the first baseline run (Step 6) or a re-run after fixes?
  5. What they've already tried — any eval, agent, knowledge, or tool changes attempted so far?

If they don't have structured results, help them organize what they have. Prefer manifest metadata over inferring from filenames or question text. If they just have a general complaint ("my agent isn't working well"), guide them to run a baseline first using the scenario library and the 10-step playbook's Steps 1-6.

Step 2: Score Interpretation

Assess readiness from the manifest's Step 4 targets and gates:

READINESS ASSESSMENT

Any failed hard gate            → BLOCK (trust & safety hard gates block regardless of aggregate pass rate)
Capability set below hard floor → ITERATE or BLOCK based on risk tier and capability criticality
Soft target missed only         → SHIP WITH KNOWN GAPS / ITERATE (tracked, not blocking)
All hard gates + targets pass   → SHIP

Setting thresholds — don't apply fixed numbers. Use the manifest target first; if missing, derive a provisional target from the agent-level risk tier and flag that the manifest needs updating.

FactorHigher Threshold When...
Criticality of errorFinancial loss, safety risk, legal exposure
ReachExternal customers or large internal population
Autonomy / blast radiusAgent can take actions or trigger downstream systems
Regulatory exposureRegulated workflow, audit requirement, compliance obligation
Data sensitivityPII, PHI, confidential, or tenant-sensitive data
Step 3: Pre-Triage Infrastructure Check

Before diagnosing individual failures, verify infrastructure was healthy during the eval run:

  • All knowledge sources accessible and fully indexed?
  • API backends and connectors returned no errors/timeouts?
  • Authentication tokens valid throughout the run?
  • Correct agent version was published and evaluated?

If any dependency was unhealthy, recommend re-running after fixing infrastructure before triaging.

Step 4: Prioritize Failures

If the user has many failures, recommend this triage order:

PriorityTriage FirstRationale
1Failed hard gates, especially trust & safety setsHighest consequence; blocks deploy regardless of aggregate score
2High-risk capability failures (accuracy, faithfulness, tool use)Direct impact on agent value; hallucination is a faithfulness/capability failure
3Lowest-scoring eval set failuresLikely systemic — fixing one pattern resolves multiple
4Recurring failures across baseline/re-runsMost diagnosable and regression-prone
5Soft-target missesImportant but non-blocking unless pattern worsens

15+ failures? Don't triage every one. Review 3-5 from the lowest-scoring eval set. If they share a root bucket/subtype pattern, fix that and re-run.

Step 5: Classify Root Cause

For each failure, work through the diagnostic questions in order. Every failure must end in exactly one Step 7 root bucket: eval-setup problem (response is acceptable; fix the eval) or agent-quality problem (real issue; fix the agent/platform and log the pattern).

TRIAGE DECISION TREE (for each failing test case)

1. Is the agent's response actually acceptable, even though it failed?
   → YES = Eval-setup problem (grader, expected value, rubric, or method is wrong)

2. Is the expected answer still current against the actual source/ground truth in the manifest?
   → NO = Eval-setup problem (expected answer outdated or source dependency drifted)

3. Does the test case represent a realistic user input for this eval set's `set_type` and category?
   → NO = Eval-setup problem (unrealistic or mis-scoped test case)

4. Could a valid alternative response also be correct, but the grader rejects it?
   → YES = Eval-setup problem (rubric/grader too rigid)

5. Is the eval method appropriate for what you're testing?
   → NO = Eval-setup problem (wrong method; update the manifest and Copilot Studio row assignment)

ALL PASS → The eval is valid. Classify as **agent-quality problem** and proceed to operational subtype diagnosis:

6. Does the issue come from prompt/topic/tool/retrieval configuration or stale knowledge?
   → YES = Agent Configuration / Knowledge Issue (agent-quality problem)

7. Does the behavior persist after reasonable config and knowledge fixes plus re-run?
   → YES = Platform Limitation (agent-quality problem; log evidence and workaround)
Step 5b: Conversation (Multi-Turn) Triage

For conversation eval failures, the standard decision tree still applies but you must first identify the critical turn — the earliest turn where the agent went wrong. Everything after a bad turn is a cascade, not independent failures.

Critical turn identification:

  1. Walk the conversation turn by turn
  2. Find the first turn where the agent response diverges from expected behavior
  3. Classify that turn using the decision tree above
  4. Mark downstream turns as "cascade — blocked by Turn N fix"

Conversation-specific failure patterns and remediations:

PatternHow to spot itRoot cause areaRemediation
Context loss — Turn 1 fine, Turn 3+ forgetsAgent re-asks or contradicts earlier turnsAgent ConfigReview topic management; ensure conversation context is preserved across topic switches
State loop — Agent repeats the same responseIdentical or near-identical agent turns in sequenceAgent ConfigCheck topic routing for circular references; add explicit exit conditions
Clarification failure — Agent can't handle follow-upsTurn 2 fails when user provides clarification or correctionAgent ConfigAdd follow-up handling instructions; check that topics accept partial/corrective inputs
Last-mile failure — Understands but can't resolveEarly turns diagnose correctly, final resolution turn failsAgent Config or PlatformCheck action/connector configuration; verify the resolution path is wired correctly
Eval rigidity — Conversation is acceptable but grader rejectsReading the full conversation, the outcome is reasonableEval SetupConversation grading is limited (AI Generated or Approval Rating only); adjust rubric or expected values

Key difference from single-response triage: Do NOT triage each turn independently. Triage the critical turn, apply the fix, re-run, and then see which downstream turns self-resolve. Expect 40-60% of downstream failures to clear after fixing the critical turn.

Two Root Buckets + Operational Subtypes
Step 7 root bucketOperational subtypeWho actsWhat it means
Eval-setup problemEval setup issueEval authorThe response is acceptable or the eval metadata/rubric/expected answer/method is wrong. Fix the eval and manifest.
Agent-quality problemAgent configuration issueAgent builderThe agent genuinely produced a bad response. Fix prompt, topics, tools, retrieval, grounding, or knowledge.
Agent-quality problemPlatform limitationPlatform team + agent ownerThe eval caught a real issue caused by platform behavior. Log evidence, workaround if possible, and track the pattern.

Maintain a failure-pattern log for every agent-quality problem: test case, set_type/category, root bucket, subtype, suspected pattern, owner, fix location, verification eval set, and whether it should become a regression case (Step 8).

Step 6: Map to Remediation

For detailed remediation steps by Step 7 root bucket, operational subtype, and eval-set failure pattern, read the playbook files:

  • Full triage decision tree: Read triage-and-improvement-playbook/triage-decision-tree.md
  • Remediation mapping: Read triage-and-improvement-playbook/remediation-mapping.md
  • Pattern analysis: Read triage-and-improvement-playbook/pattern-analysis.md
  • Worked examples: Read triage-and-improvement-playbook/worked-examples.md
Quick Remediation Reference

Eval-setup fixes:

Sub-TypeFix
Outdated expected answerUpdate expected value to match current source content
Overly rigid graderSwitch to Compare Meaning, or broaden keyword set
Unrealistic test caseRewrite input using actual user language
Wrong eval methodChange method to match the eval-set purpose and evidence type
Grader error/biasReview rubric, add examples, consider deterministic method

Agent-quality fixes — agent configuration / knowledge / tools:

Failure patternCommon Fix
Factual accuracy (wrong source)Review knowledge source config, verify indexing, check vocabulary match
Factual accuracy (wrong extraction)Add extraction guidance to system prompt
Hallucination (faithfulness capability failure)Improve retrieval/chunking first; add instruction: "Only answer from knowledge sources. If unavailable, say so."
Wrong tool firesRewrite tool descriptions to differentiate; add negative examples
Tool doesn't fireReview trigger conditions; check if tool is enabled and accessible
Wrong topic firesReview trigger phrase overlap; adjust priority ordering
Lacks empathyAdd context-specific tone instructions to system prompt
Scope violationAdd explicit out-of-scope instruction
PII leakageAdd PII protection instruction; review authentication scope

Agent-quality fixes — platform limitation response:

  • Document the limitation with evidence
  • Implement workaround where possible
  • Adjust eval thresholds to account for known platform behavior
  • File with platform team with reproduction steps
Step 7: Triage Rationale (teach the WHY)

Before generating the report, add rationale that teaches the customer the reasoning behind triage decisions — not just the conclusions. For each of these, use the actual eval data from this triage:

  1. Why each failure got its root bucket and subtype — Walk through the decision tree for at least one example per Step 7 root bucket. E.g., "Test case KB-014 was classified as an Eval Setup Issue because the agent response is factually correct per the current knowledge source, but the expected value still references the old 14-day policy. The agent is right; the eval is stale."

  2. Why the remediation targets config vs. content vs. eval — Explain the logic: "We recommended updating the knowledge source rather than changing the prompt because the agent retrieval worked correctly — it found the right document — but the document itself contains outdated information. A prompt change would mask the real problem."

  3. Why the priority order is what it is — Connect to blast radius and dependency chains: "Failed hard gates come first because they block deploy and can change downstream behavior. Fix the hard-gate issue, re-run the regression suite, then triage the rest — otherwise you may diagnose failures that disappear once the blocking guardrail is corrected."

  4. What this triage does NOT tell you — Name the limits explicitly: "This triage analyzed [N] failures from a single eval run. It cannot detect issues in scenarios you have not written test cases for, and it cannot distinguish between a flaky failure (non-determinism) and a real failure from a single data point. If a failure is borderline, re-run before investing in a fix."

Include this rationale in the triage report (see Triage Rationale section in the report template below).

Show full SKILL.md (899 more words)Show less
Step 8: Generate Triage Report

Output a structured triage report:

markdown
# Triage Report: [Agent Name] — [Date]

## Score Summary
| Eval Set | Set Type / Category | Pass Rate | Target | Gate | Status |
|----------|---------------------|-----------|--------|------|--------|
| ... | capability / faithfulness | ... | ... | hard/soft | PASS/BLOCK/ITERATE |

## Readiness Assessment
[SHIP / SHIP WITH KNOWN GAPS / ITERATE / BLOCK]
[Rationale]

## Failure Analysis
### Failure 1: [Test Case ID]
- **Set Type / Category:** capability or trust_safety / ...
- **Eval-set focus:** ...
- **Sample Input:** ...
- **Expected:** ...
- **Actual:** ...
- **Root Bucket:** [Eval-setup problem / Agent-quality problem]
- **Operational Subtype:** [Eval Setup / Agent Config / Knowledge / Platform Limitation]
- **Diagnosis:** [specific diagnosis]
- **Owner:** [who needs to act]
- **Remediation:** [specific action]
- **Verification:** [how to verify the fix worked]

[Repeat for each triaged failure]

## Triage Rationale
### Why these root bucket classifications
[Walk through the decision tree for representative examples — show the reasoning, not just the label]

### Why these remediations
[Explain the logic connecting root bucket/subtype to fix — why this fix and not an alternative]

### Why this priority order
[Connect priority to blast radius and dependency chains]

### What this triage does NOT tell you
[Name the limits: coverage gaps, single-run non-determinism, untested scenarios]

## Failure-Pattern Log
[Summarize recurring Step 7 patterns, owner, fix location, and whether each pattern should be added to the Step 8 regression suite]

## Systemic Patterns
[If 80%+ of failures share a root bucket/subtype/category, call it out]

## Action Items
| # | Action | Owner | Priority | Verification |
|---|--------|-------|----------|-------------|
| 1 | ... | ... | ... | Re-run [eval set] |

## Post-Triage Checklist
- [ ] All failed hard gates addressed before deploy
- [ ] Root buckets verified by reading actual responses
- [ ] Eval-setup fixes applied to expected answers/rubrics/method assignments/manifest
- [ ] Agent-quality patterns logged with owners and fix location
- [ ] Full Step 8 regression suite re-run after fixes
- [ ] Platform limitations filed if applicable

## Human Review Required
[Include human review checkpoints table — see Human Review Checkpoints section below]
Post-Triage Verification

After fixes are applied:

  • Scores flat after fix? → Wrong root bucket/subtype, re-triage
  • One score up, another down? → Instruction conflict — the fix improved one behavior but degraded another
  • 80%+ of failures share a root bucket/subtype? → Systemic issue — fix the category, not individual test cases

Non-Determinism Handling

LLM-based agents and graders produce variable outputs:

  • Establish baselines: Run 3+ times before treating any score as the Step 6 baseline. Use the average and record agent version + timestamp.
  • Normal variance: +/-5% between runs is expected. Investigate if >10%.
  • Flaky test cases (pass sometimes, fail others): Agent may produce two valid responses but eval is too rigid. Investigate whether to broaden the expected value.
  • Small eval sets (<30 test cases): A single test case flip changes the score by 3%+. Don't over-interpret.

Step 9 Optimization Loop: Production Signals

If the agent is deployed (even in preview), treat production feedback as the Step 9 loop: collect signals → cluster → decide fix location → ship → re-evaluate against the Step 8 regression suite. Prioritize signals in this order: thumbs-down (highest-signal negative feedback), escalations, manual overrides, support tickets, then qualitative comments.

  • High thumbs-down on a topic where eval passes: Coverage gap. Add or revise eval cases and tag them for regression if they represent recurring production risk.
  • Thumbs-down clustering after a config change: Possible regression. Re-run the Step 8 regression suite and add a case if the suite missed it.
  • Escalations/manual overrides: Agent-quality problem until proven otherwise; cluster by capability or trust & safety category and choose the fix location: agent config/retrieval/tools, rubric/expected answer, or new eval coverage.
  • Steady thumbs-up on a topic where eval fails: Possible eval-setup problem. Review the actual responses before weakening a gate.

Production signals are not verdicts by themselves. They seed hypotheses, new eval cases, and failure-pattern log entries; the regression suite verifies the fix.

Human Review Checkpoints

Before acting on the triage report, review these checkpoints. Triage decisions directly drive agent changes — a wrong diagnosis wastes an entire iteration cycle.

#CheckpointWhy it matters
1Verify Step 7 root buckets yourself — For each failure classified as an eval-setup problem, read the agent actual response. Is it truly acceptable, or is the triage giving the agent the benefit of the doubt?Misclassifying agent-quality problems as eval setup means real problems get ignored. The two-bucket distinction requires judgment, not score-only automation.
2Confirm systemic pattern diagnoses before applying systemic fixes — If the report says 80%+ failures share a root bucket/subtype, verify by reading the actual responses. Similar symptoms can have different causes.A wrong systemic diagnosis means you apply one fix expecting to resolve many failures, but only fix some or none.
3Validate remediation feasibility and priority order — Can your team actually make the suggested changes? Is the priority order right for your timeline and constraints?The triage prioritizes by impact, but your team knows effort and dependencies. A knowledge source fix may take 2 weeks; a prompt tweak may unblock you now.
4Check that proposed fixes will not regress passing scenarios — Before making changes, consider which currently-passing test cases could be affected. Prompt changes especially have ripple effects.Fixing 3 failures while introducing 5 new ones is a net loss. Plan to re-run the full suite after any agent configuration change.
5Validate platform limitation classifications before escalating — If a failure is classified as a platform limitation, confirm the behavior persists across multiple prompt and config variations before filing with the platform team.Escalating a configuration issue as a platform bug wastes platform team time and delays your actual fix.
6Review manifest targets and gates against your actual risk tier — Does SHIP/ITERATE/BLOCK honor hard gates, soft targets, and the five risk factors?Only your team knows your real risk tolerance. A soft-target miss may be acceptable for an internal helper but a failed hard trust & safety gate blocks deploy.

Include this table in the triage report output. Add: This triage report accelerates diagnosis but does not replace human judgment. Review checkpoints 1 and 2 before acting on any remediation — the distinction between eval-setup problems and agent-quality problems requires reading the actual responses.

Data Retention Warning

Copilot Studio deletes test run results after 89 days. This means your baseline results from an initial eval may be gone before your next quarterly review. After every triage cycle:

  1. Export the results CSV immediately (Test set → Export results)
  2. Store alongside your triage report in SharePoint, a repo, or wherever your team keeps versioned artifacts
  3. Tag with agent version and date so future comparisons are possible

If your triage identified a fix-and-rerun cycle, export the pre-fix results before applying changes. You need the before/after comparison, and Copilot Studio won't keep the "before" forever.

Cross-Reference

This skill uses skills/eval-guide/playbook.md as the methodology spine. It also works alongside the AI Agent Evaluation Scenario Library (github.com/microsoft/ai-agent-eval-scenario-library), which defines supporting scenario patterns and quality dimensions, and the Triage & Improvement Playbook (github.com/microsoft/triage-and-improvement-playbook), which provides supporting diagnostic frameworks for Step 7.

After triage, if you need to...Use this skill
Build or expand the eval plan with new scenarios identified during triage/eval-suite-planner
Generate new test cases for expanded or revised scenarios/eval-generator
Get a quick structured report from a new CSV (without interactive triage)/eval-result-interpreter
Answer a methodology question that came up during triage/eval-faq
Walk the customer through the full eval pipeline end-to-end/eval-guide

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/eval-triage-and-improvement of microsoft/eval-guide.

Open the folder on GitHubat commit 7a22a89

Compare with similar skills

Eval Triage And Improvement next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Triage And Improvement compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Triage And Improvement this skillmicrosoft/eval-guide138—~5.9kAutomated safety check: PassMIT
Blue Teamgaasher/Agent-Loop-Skills174—~3.6kAutomated safety check: PassMIT
Pester Failure AnalysisPowerShell/PowerShell56k—~5.1kAutomated safety check: PassMIT
Skill Testdatabricks-solutions/ai-dev-kit1.9k—~1.9kAutomated safety check: PassCustom licence
Code SolvingHoangTheQuyen/think-better122—~3.7kAutomated safety check: PassMIT
Ue Test AuthoringJasonMa0012/MooaToon750—~2.1kAutomated safety check: NotesCustom licence

Similar skills

  • Blue Team

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has concrete failing cases in code or a guardrail/classifier/filter/prompt/API they own — a red-team failure catalogue OR a CI/CD test-failure report (failing…

    174 GitHub stars~3.6k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Pester Failure Analysis

    PowerShell/PowerShell

    Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations.

    56k GitHub stars~5.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Skill Test

    databricks-solutions/ai-dev-kit

    Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit.

    1.9k GitHub stars~1.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Code Solving

    HoangTheQuyen/think-better

    Structured coding workflow for non-trivial code work: debug, build features, refactor, optimize, migrate and review code through 7 steps with evidence-based quality gates.

    122 GitHub stars~3.7k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Ue Test Authoring

    JasonMa0012/MooaToon

    A skill your agent uses when writing or modifying UE automated tests (Automation, CQTest, Functional, Gauntlet, LowLevel) with Rider MCP available.

    750 GitHub stars~2.1k tokensUpdated 21 days ago
    Testing & QAAuto-check: notes
  • Diagnose

    avibebuilder/claude-prime

    Investigate unexpected behavior and mysterious bugs. An agent skill from avibebuilder/claude-prime.

    120 GitHub stars~1.2k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed

More from microsoft/eval-guide

  • Eval Guide

    microsoft/eval-guide

    Official

    Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

    138 GitHub stars~22k tokensUpdated 3 mo ago
    Auto-check: warnings
  • Eval Faq

    microsoft/eval-guide

    Official

    Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage &…

    138 GitHub stars~10k tokensUpdated 3 mo ago
    Auto-check passed
  • Eval Generator

    microsoft/eval-guide

    Official

    Generate standalone — turns the populated Eval Suite Planning workbook (output of /eval-suite-planner) into concrete capability eval sets and trust & safety eval sets.

    138 GitHub stars~7.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Eval Result Interpreter

    microsoft/eval-guide

    Official

    Analyzes Copilot Studio evaluation results using Practical Guidance on Agent Evaluation's 10-step playbook (Steps 6, 7, and 9) plus Microsoft's triage diagnostics.

    138 GitHub stars~9.9k tokensUpdated 3 mo ago
    Auto-check passed
  • Eval Suite Planner

    microsoft/eval-guide

    Official

    Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.

    138 GitHub stars~2.3k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Eval Triage And Improvement

What does Eval Triage And Improvement do?

A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps…. Eval Triage And Improvement is an agent skill from microsoft/eval-guide, published by the product's own GitHub organization. Use this skill when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps, or analyze patterns to improve their agent.

When should I use Eval Triage And Improvement?

Eval Triage And Improvement fits situations like: the users Copilot Studio agent evaluations have come back and they need to interpret scores; diagnose root causes of underperforming test cases; find remediation steps; analyze patterns to improve their agent.

How do I install Eval Triage And Improvement in Claude Code?

Run `npx skills add microsoft/eval-guide --skill eval-triage-and-improvement -a claude-code`. Or copy the skill folder (skills/eval-triage-and-improvement in microsoft/eval-guide) into .claude/skills/eval-triage-and-improvement in your project. Claude Code loads it when a task matches its description.

How do I install Eval Triage And Improvement in Codex?

Run `npx skills add microsoft/eval-guide --skill eval-triage-and-improvement -a codex`. Or copy the skill folder (skills/eval-triage-and-improvement in microsoft/eval-guide) into .agents/skills/eval-triage-and-improvement in your project. Codex loads it when a task matches its description.

Can I use Eval Triage And Improvement in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/eval-guide --skill eval-triage-and-improvement -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-triage-and-improvement, .gemini/skills/eval-triage-and-improvement, .github/skills/eval-triage-and-improvement and .opencode/skills/eval-triage-and-improvement in your project.

What does Eval Triage And Improvement need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Triage And Improvement is instructions for the agent only.

Does Eval Triage And Improvement access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Triage And Improvement safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Triage And Improvement use?

Eval Triage And Improvement is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Triage And Improvement use?

About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Triage And Improvement?

Skills that share tags, products or a category with Eval Triage And Improvement: Blue Team (gaasher/Agent-Loop-Skills, 174 stars), Pester Failure Analysis (PowerShell/PowerShell, 56k stars), Skill Test (databricks-solutions/ai-dev-kit, 1.9k stars) and Code Solving (HoangTheQuyen/think-better, 122 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Triage And Improvement?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/eval-guide, which has 138 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on June 24, 2026.

Source: microsoft/eval-guide on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.