Official agent skill

Evaluate PR Tests

by dotnet in dotnet/maui

Reviews the tests added in a pull request for fix coverage, quality, edge cases and test type, and recommends lighter test types where they would do.

OfficialMITAuto-check passedTesting & QA

Install Evaluate PR Tests

skills CLI
$ npx skills add dotnet/maui --skill evaluate-pr-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dotnet/maui evaluate-pr-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dotnet/maui.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/evaluate-pr-tests .claude/skills/evaluate-pr-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate-pr-tests
GitHub stars
23k
Token cost
~2.9k tokens
SKILL.md length
1,055 words
Files
3 (incl. scripts)
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Reviews the tests added in a pull request for fix coverage, quality, edge cases and test type, and recommends lighter test types where they would do.

  • Works in 12 steps: Gather Automated Context → Understand the Fix → Evaluate the Tests → …
  • Reviewing whether the tests in a PR actually cover the fix
  • SKILL.md covers When to Use, Quick Start, Workflow and Evaluation Criteria, plus 3 more sections
  • Runs PowerShell scripts from its folder; calls pwsh

What it does

The skill reviews the tests added in a pull request and produces a structured report with actionable findings. It first runs `Gather-TestContext.ps1`, which writes a context report covering file categorization (fix files versus test files), convention checks for naming, attributes and anti-patterns, AutomationId consistency between the host app and the tests, similar existing tests and platform scope. The agent then reads the fix to learn what changed, why and which edge cases exist.

Each test file is judged against a list of criteria, starting with fix coverage: does the test exercise the changed code path, would it fail if the fix were reverted, and does it assert the behavior that was broken. Every criterion gets a pass, concern or fail verdict with an explanation. The skill also pushes toward lighter test types, preferring unit tests over device tests over UI tests, and a PR with no tests gets a fail verdict on coverage. The excerpt ends within the criteria list.

When your agent uses it

  • Reviewing whether the tests in a PR actually cover the fix
  • Checking if a lighter test type could replace a UI test
  • Assessing test quality before merging a pull request

Example prompts

  • “Evaluate the tests in this PR and tell me whether they would fail if the fix were reverted.”
  • “Are these tests good enough? Check whether a unit test could replace the UI test.”
  • “Assess test coverage for the PR before I merge it.”

Requirements

  • git, PowerShell and the gh CLI for PR context
  • Compatibility (from SKILL.md): Requires git, PowerShell, and gh CLI for PR context.

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Gather Automated Context
  2. Understand the Fix
  3. Evaluate the Tests
  4. Produce the Report
  5. Fix Coverage
  6. Edge Cases & Gaps
  7. Test Type Appropriateness
  8. Convention Compliance
  9. Flakiness Risk
  10. Duplicate Coverage
  11. Platform Scope
  12. Assertion Quality

What it can do on your machine

Read from SKILL.md and the folder at commit 7d38fd0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (PowerShell), which the agent can run.

    Shell commands in SKILL.md call:

    • pwsh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires git, PowerShell, and gh CLI for PR context.

    From compatibility in the SKILL.md frontmatter.

Context cost

Evaluate PR Tests loads about 2.9k tokens when it runs. Until then it costs about 107 tokens; SKILL.md has 1,055 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from dotnet/maui at commit 7d38fd0, republished under its MIT licence (© dotnet). 1,055 words, ~2,949 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate-pr-tests/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evaluate-pr-tests
description
Evaluates tests added in a PR for coverage, quality, edge cases, and test type appropriateness. Checks if tests cover the fix, finds gaps, and recommends lighter test types when possible. Prefer unit tests over device tests over UI tests. Triggers on: 'evaluate tests in PR', 'review test quality', 'are these tests good enough', 'check test coverage', 'is this test adequate', 'assess test coverage for PR'.
compatibility
Requires git, PowerShell, and gh CLI for PR context.
metadata.author
dotnet-maui
metadata.version
1.0

Evaluate PR Tests

Evaluates the quality, coverage, and appropriateness of tests added in a PR. Produces a structured report with actionable findings.

When to Use

  • ✅ PR has tests and you want to evaluate their quality
  • ⚠️ PR has no test files -- output a ❌ Fix Coverage verdict noting no tests were added; skip remaining criteria
  • ✅ Reviewing whether tests adequately cover the fix
  • ✅ Checking if a lighter test type could be used instead
  • ✅ Before merging a PR, as part of review

Quick Start

bash
# Auto-detect PR and base branch
pwsh .github/skills/evaluate-pr-tests/scripts/Gather-TestContext.ps1

# With explicit base branch
pwsh .github/skills/evaluate-pr-tests/scripts/Gather-TestContext.ps1 -BaseBranch "origin/main"

Workflow

Step 1: Gather Automated Context

Run the script to get file categorization, convention checks, and anti-pattern detection:

bash
pwsh .github/skills/evaluate-pr-tests/scripts/Gather-TestContext.ps1

This produces a report at CustomAgentLogsTmp/TestEvaluation/context.md with:

  • File categorization (fix files vs test files by type)
  • Convention compliance checks (naming, attributes, anti-patterns)
  • AutomationId consistency (HostApp ↔ test)
  • Existing similar tests
  • Platform scope analysis
Step 2: Understand the Fix

Read the fix files to understand:

  • What changed — which code paths were modified
  • Why it changed — the bug being fixed (from PR description or linked issue)
  • Edge cases — what boundary conditions exist in the changed code
Step 3: Evaluate the Tests

Read each test file and evaluate against all criteria below. For each criterion, provide a verdict (✅ Pass, ⚠️ Concern, ❌ Fail) with explanation.

Step 4: Produce the Report

Output a structured evaluation report (see Output Format below).


Evaluation Criteria

1. Fix Coverage

Question: Does the test exercise the actual code paths changed by the fix?

How to check:

  • Trace the test's actions through the code to the fix location
  • Would the test fail if the fix were reverted?
  • Does the test assert on the specific behavior that was broken?

Red flags:

  • Test only checks that a page loads (doesn't exercise the fix)
  • Test asserts on a different property/behavior than what was fixed
  • Test interacts with the control but doesn't trigger the buggy code path

Example — Good:

csharp
// Fix: CollectionView.SelectedItem setter now clears selection when set to null
// Test: Sets SelectedItem to null and verifies selection is cleared
App.Tap("SelectItem");
App.Tap("ClearSelection");  // Sets SelectedItem = null
var text = App.FindElement("SelectionStatus").GetText();
Assert.That(text, Is.EqualTo("None"));  // Directly tests the fix

Example — Bad:

csharp
// Fix: CollectionView.SelectedItem setter
// Test: Just checks CollectionView renders (doesn't test selection clearing)
App.WaitForElement("MyCollectionView");
Assert.That(true);  // Proves nothing about the fix
2. Edge Cases & Gaps

Question: Does the test cover boundary conditions, or only the happy path?

Check for these common gaps:

Gap TypeWhat to Look For
Null/emptyDoes the fix handle null? Is it tested?
Boundary valuesMin, max, zero, negative, very large
Repeated actionsDoes calling the action twice cause issues?
Platform-specificDoes the bug only occur on certain platforms?
Async/timingDoes the fix involve async code? Race conditions?
State transitionsDoes the test cover before→after state changes?
Error pathsWhat happens when the operation fails?
Combination effectsDoes the fix interact with other properties/features?

How to suggest missing edge cases:

  • Read the fix code and identify every conditional branch
  • For each branch, check if the test covers it
  • Look for if (x == null), if (x <= 0), try/catch blocks
  • Consider: "What inputs would make this fix NOT work?"
3. Test Type Appropriateness

Question: Is this the lightest test type that can verify the fix?

Preference order (lightest → heaviest):

PriorityTypeWhen AppropriateProject
⭐ 1stUnit TestPure logic, property changes, data transformations, binding behavior, event wiring*.UnitTests.csproj
⭐ 1stXAML TestXAML parsing, XamlC compilation, source generation, markup extensionsControls.Xaml.UnitTests
⭐⭐ 2ndDevice TestPlatform-specific rendering, native API interaction, handler mapping*.DeviceTests.csproj
⭐⭐⭐ 3rdUI TestUser interaction flows, visual layout, screenshot comparison, end-to-end scenariosTestCases.Shared.Tests

Decision tree:

Does the test need to interact with visual UI elements?
  YES → Is it checking visual layout/appearance?
    YES → UI test (VerifyScreenshot) ✅
    NO  → Could the interaction be tested via handler/control API?
      YES → Device test ⭐⭐
      NO  → UI test ✅
  NO  → Does it need a platform/native context?
    YES → Device test ⭐⭐
    NO  → Does it test XAML parsing/compilation?
      YES → XAML test ⭐
      NO  → Unit test ⭐

Common "could be lighter" patterns:

Current Test DoesCould Be InsteadWhy
UI test: sets property, checks label textUnit testProperty logic doesn't need UI
UI test: verifies event firesUnit testEvent wiring is testable in isolation
UI test: checks control doesn't crashDevice testDon't need Appium for crash testing
UI test: validates XAML bindingXAML testBinding resolution is compile-time
Device test: checks property defaultUnit testDefaults don't need platform context
Show full SKILL.md (460 more words)Show less
4. Convention Compliance

Automated by the script. Review the script output for:

UI Tests:

  • File naming: IssueXXXXX.cs
  • [Issue()] attribute on HostApp page
  • [Category()] attribute — exactly ONE per test class (on the class or method, not both)
  • _IssuesUITest base class
  • WaitForElement before interactions
  • No Task.Delay/Thread.Sleep
  • No inline #if ANDROID/#if IOS
  • No obsolete APIs (Application.MainPage, Frame, Device.BeginInvokeOnMainThread)
  • UITestEntry/UITestEditor for screenshot tests

Unit Tests:

  • [Fact] or [Theory] attributes (xUnit)

XAML Tests:

  • [Test] with [Values] XamlInflator parameter
  • Issue naming: MauiXXXXX
5. Flakiness Risk

Question: Is this test likely to be flaky in CI?

Risk FactorDetectionMitigation
Arbitrary delaysTask.Delay, Thread.SleepUse WaitForElement, retryTimeout
Missing waitsApp.Tap without prior WaitForElementAdd explicit waits
Screenshot timingVerifyScreenshot() without retryTimeoutAdd retryTimeout: TimeSpan.FromSeconds(2)
Cursor blinkEntry/Editor in screenshot testUse UITestEntry/UITestEditor
External URLsWebView loading remote contentUse mock URLs or local content
Animation timingVisual check after animationUse retryTimeout
Global stateTest modifies Application.CurrentEnsure cleanup in teardown
6. Duplicate Coverage

Question: Does a similar test already exist?

Check the "Existing Similar Tests" section of the script output. If similar tests exist:

  • Is the new test covering a different scenario? → OK
  • Is the new test redundant? → Flag as concern
  • Could the new test be merged with an existing one? → Suggest consolidation
7. Platform Scope

Question: Does the test run on all platforms affected by the fix?

Check the "Platform Scope Analysis" from the script:

  • Cross-platform fix → tests should run on all platforms
  • Platform-specific fix → test on that platform is sufficient
  • Fix affects iOS + MacCatalyst → both should be tested (.ios.cs compiles for both)
8. Assertion Quality

Question: Are the assertions specific enough to catch regressions?

Assertion QualityExampleVerdict
✅ SpecificAssert.That(label.Text, Is.EqualTo("Expected Value"))Catches regression
⚠️ VagueAssert.That(label.Text, Is.Not.Null)Too permissive
❌ MeaninglessAssert.That(true) or no assertionProves nothing
✅ PositionalAssert.That(rect.Y, Is.GreaterThan(safeAreaTop))Specific to layout fix
⚠️ BrittleAssert.That(rect.Y, Is.EqualTo(47))Magic number, will break
9. Fix-Test Alignment

Question: Do the files changed by the fix align with what the test exercises?

  • Map the fix files to the controls/features they affect
  • Map the test to the controls/features it exercises
  • Flag if test exercises a different control than the fix changes
  • Flag if test only covers one platform when fix touches multiple

Red flags:

  • Test class is named Issue12345 for a fix in CollectionView but only exercises Label rendering
  • Fix changes Shell.cs but test only navigates a ContentPage

Output Format

Produce the evaluation report in this format:

markdown
## PR Test Evaluation Report

**PR:** #XXXXX — [Title]
**Test files evaluated:** [count]
**Fix files:** [count]

---

### Overall Verdict

[One of: ✅ Tests are adequate | ⚠️ Tests need improvement | ❌ Tests are insufficient]

[1-2 sentence summary of the most important finding]

---

### 1. Fix Coverage — [✅/⚠️/❌]

[Does the test exercise the code paths changed by the fix?]

### 2. Edge Cases & Gaps — [✅/⚠️/❌]

**Covered:**
- [edge case 1]
- [edge case 2]

**Missing:**
- [gap 1 — describe what should be tested and why]
- [gap 2]

### 3. Test Type Appropriateness — [✅/⚠️/❌]

**Current:** [UI Test / Device Test / Unit Test / XAML Test]
**Recommendation:** [Same / Could be lighter — explain why]

### 4. Convention Compliance — [✅/⚠️/❌]

[Summary from automated checks — list only issues found]

### 5. Flakiness Risk — [✅ Low / ⚠️ Medium / ❌ High]

[Specific risk factors identified]

### 6. Duplicate Coverage — [✅ No duplicates / ⚠️ Potential overlap]

[Similar existing tests found, if any]

### 7. Platform Scope — [✅/⚠️/❌]

[Does test coverage match the platforms affected by the fix?]

### 8. Assertion Quality — [✅/⚠️/❌]

[Are assertions specific enough to catch the actual bug?]

### 9. Fix-Test Alignment — [✅/⚠️/❌]

[Do the test and fix target the same code paths?]

---

### Recommendations

1. [Most important actionable recommendation]
2. [Second recommendation]
3. [...]

Output Files

FileDescription
CustomAgentLogsTmp/TestEvaluation/context.mdAutomated context report from script

Troubleshooting

ProblemCauseSolution
No changed files detectedWrong base branchUse -BaseBranch explicitly
No fix files detectedAll changes are testsExpected for test-only PRs
AutomationId mismatchHostApp and test out of syncUpdate one to match the other
Convention check false positiveScript regex too broadIgnore and note in report

© dotnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in .github/skills/evaluate-pr-tests of dotnet/maui.

  • SKILL.md
  • scripts/Gather-TestContext.ps1
  • tests/eval.vally.yaml

Open the folder on GitHubat commit 7d38fd0

Compare with similar skills

Evaluate PR Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate PR Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate PR Tests this skilldotnet/maui23k—~2.9kAutomated safety check: PassMIT
Differential Security Reviewtrailofbits/skills7.4k—~1.8kAutomated safety check: NotesCC-BY-SA-4.0
Code Review WorkflowResgrid/Core229—~2.1kAutomated safety check: PassApache-2.0
Reviewatelier-fashion/adlc-toolkit171—~1.9kAutomated safety check: PassMIT
Code Reviewpolyipseity/obsidian-terminal948—~1.6kAutomated safety check: PassAGPL-3.0
Core Components Code Reviewcore-ds/core-components137—~5.4kAutomated safety check: PassMIT

Similar skills

  • Official

    Reviews a pull request, commit or diff for security problems, using git history, caller counts and test coverage, and writes a markdown report.

    7.4k GitHub stars~1.8k tokensUpdated 5 days ago
    SecurityAuto-check: notes
  • Structured code review workflow for .NET projects using Roslyn MCP tools.

    229 GitHub stars~2.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Review

    atelier-fashion/adlc-toolkit

    Multi-agent code review covering correctness, quality, architecture, test coverage, and security

    171 GitHub stars~1.9k tokensUpdated 8 days ago
    Testing & QAAuto-check passed
  • Code Review

    polyipseity/obsidian-terminal

    A skill your agent uses when reviewing PRs, code changes, or conducting code audits in obsidian-terminal.

    948 GitHub stars~1.6k tokensUpdated 4 days ago
    DevelopmentAuto-check passed
  • Core Components Code Review

    core-ds/core-components

    Review a Pull Request or diff in the @alfalab/core-components UI library — correctness bugs, public API/breaking changes, accessibility, keyboard/focus/pointer interaction, component states…

    137 GitHub stars~5.4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • PR Deep Verification

    QwenLM/qwen-code

    Runs a sandboxed, evidence-based check of one qwen-code pull request, proving its main change against the base build and writing a report with a machine-readable verdict.

    28k GitHub stars~18k tokensUpdated today
    DevelopmentAuto-check passed

More from dotnet/maui

All 27 skills in this repo
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Official

    Produces evidence-backed ship-readiness verdicts for .NET MAUI Servicing Releases and Previews, and drafts public-safe release handoff pages from the result.

    23k GitHub stars~15k tokensUpdated today
    Auto-check passed
  • Official

    Interprets pinned managed benchmark evidence for a dotnet/maui pull request and writes a narrative for the performance review workflow, without running or publishing anything.

    23k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • PR Finalize

    dotnet/maui

    Official

    Checks that a pull request's title and description match its implementation and reviews the code for best practices before merge, without posting anything.

    23k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Official

    Adds MAUI-specific guardrails on top of the maestro-cli skill and Maestro MCP tools for darc, BAR, and channel or feed lookups in dotnet/maui.

    23k GitHub stars~10k tokensUpdated today
    Auto-check passed
  • Official

    Adds dotnet/maui-specific context for investigating failing PR checks and broken nightly builds: pipelines, Helix logs, binlogs and merge-readiness verdicts.

    23k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Evaluate PR Tests

What does Evaluate PR Tests do?

Reviews the tests added in a pull request for fix coverage, quality, edge cases and test type, and recommends lighter test types where they would do. The skill reviews the tests added in a pull request and produces a structured report with actionable findings.ps1`, which writes a context report covering file categorization (fix files versus test files), convention checks for naming, attributes and anti-patterns, AutomationId consistency between the host app and the tests, similar existing tests and platform scope.

When should I use Evaluate PR Tests?

Evaluate PR Tests fits situations like: reviewing whether the tests in a PR actually cover the fix; checking if a lighter test type could replace a UI test; assessing test quality before merging a pull request.

How do I install Evaluate PR Tests in Claude Code?

Run `npx skills add dotnet/maui --skill evaluate-pr-tests -a claude-code`. Or copy the skill folder (.github/skills/evaluate-pr-tests in dotnet/maui) into .claude/skills/evaluate-pr-tests in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate PR Tests in Codex?

Run `npx skills add dotnet/maui --skill evaluate-pr-tests -a codex`. Or copy the skill folder (.github/skills/evaluate-pr-tests in dotnet/maui) into .agents/skills/evaluate-pr-tests in your project. Codex loads it when a task matches its description.

Can I use Evaluate PR Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dotnet/maui --skill evaluate-pr-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-pr-tests, .gemini/skills/evaluate-pr-tests, .github/skills/evaluate-pr-tests and .opencode/skills/evaluate-pr-tests in your project.

What does Evaluate PR Tests need to run?

Going by SKILL.md and its folder, Evaluate PR Tests needs PowerShell for the scripts in its folder and the command-line tools its instructions call (pwsh). Our summary lists: git, PowerShell and the gh CLI for PR context. Compatibility (from SKILL.md): Requires git, PowerShell, and gh CLI for PR context..

Does Evaluate PR Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluate PR Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evaluate PR Tests use?

Evaluate PR Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate PR Tests use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluate PR Tests?

Skills that share tags, products or a category with Evaluate PR Tests: Differential Security Review (trailofbits/skills, 7.4k stars), Code Review Workflow (Resgrid/Core, 229 stars), Review (atelier-fashion/adlc-toolkit, 171 stars) and Code Review (polyipseity/obsidian-terminal, 948 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate PR Tests?

dotnet (a GitHub organization, an official publisher) maintains it in dotnet/maui, which has 23,322 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 7, 2026.

Source: dotnet/maui on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.