Differential Security Review
trailofbits/skills
Reviews a pull request, commit or diff for security problems, using git history, caller counts and test coverage, and writes a markdown report.
Reviews the tests added in a pull request for fix coverage, quality, edge cases and test type, and recommends lighter test types where they would do.
$ npx skills add dotnet/maui --skill evaluate-pr-tests -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install dotnet/maui evaluate-pr-tests --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/dotnet/maui.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/evaluate-pr-tests .claude/skills/evaluate-pr-tests && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluate-pr-tests" agent skill from https://github.com/dotnet/maui/tree/main/.github/skills/evaluate-pr-tests into .claude/skills/evaluate-pr-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-pr-tests", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/dotnet/maui/tree/main/.github/skills/evaluate-pr-testsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add dotnet/maui --skill evaluate-pr-tests -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install dotnet/maui evaluate-pr-tests --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dotnet/maui.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.github/skills/evaluate-pr-tests .agents/skills/evaluate-pr-tests && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluate-pr-tests" agent skill from https://github.com/dotnet/maui/tree/main/.github/skills/evaluate-pr-tests into .agents/skills/evaluate-pr-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-pr-tests", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dotnet/maui --skill evaluate-pr-tests -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install dotnet/maui evaluate-pr-tests --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dotnet/maui.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.github/skills/evaluate-pr-tests .cursor/skills/evaluate-pr-tests && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluate-pr-tests" agent skill from https://github.com/dotnet/maui/tree/main/.github/skills/evaluate-pr-tests into .cursor/skills/evaluate-pr-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-pr-tests", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/dotnet/maui.git --path .github/skills/evaluate-pr-tests--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add dotnet/maui --skill evaluate-pr-tests -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install dotnet/maui evaluate-pr-tests --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dotnet/maui.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.github/skills/evaluate-pr-tests .gemini/skills/evaluate-pr-tests && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluate-pr-tests" agent skill from https://github.com/dotnet/maui/tree/main/.github/skills/evaluate-pr-tests into .gemini/skills/evaluate-pr-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-pr-tests", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install dotnet/maui evaluate-pr-testsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add dotnet/maui --skill evaluate-pr-tests -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/dotnet/maui.git skills-src && mkdir -p .github/skills && cp -r skills-src/.github/skills/evaluate-pr-tests .github/skills/evaluate-pr-tests && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluate-pr-tests" agent skill from https://github.com/dotnet/maui/tree/main/.github/skills/evaluate-pr-tests into .github/skills/evaluate-pr-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-pr-tests", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dotnet/maui --skill evaluate-pr-tests -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install dotnet/maui evaluate-pr-tests --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dotnet/maui.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.github/skills/evaluate-pr-tests .opencode/skills/evaluate-pr-tests && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluate-pr-tests" agent skill from https://github.com/dotnet/maui/tree/main/.github/skills/evaluate-pr-tests into .opencode/skills/evaluate-pr-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-pr-tests", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluate-pr-testsReviews the tests added in a pull request for fix coverage, quality, edge cases and test type, and recommends lighter test types where they would do.
The skill reviews the tests added in a pull request and produces a structured report with actionable findings. It first runs `Gather-TestContext.ps1`, which writes a context report covering file categorization (fix files versus test files), convention checks for naming, attributes and anti-patterns, AutomationId consistency between the host app and the tests, similar existing tests and platform scope. The agent then reads the fix to learn what changed, why and which edge cases exist.
Each test file is judged against a list of criteria, starting with fix coverage: does the test exercise the changed code path, would it fail if the fix were reverted, and does it assert the behavior that was broken. Every criterion gets a pass, concern or fail verdict with an explanation. The skill also pushes toward lighter test types, preferring unit tests over device tests over UI tests, and a PR with no tests gets a fail verdict on coverage. The excerpt ends within the criteria list.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 7d38fd0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (PowerShell), which the agent can run.
Shell commands in SKILL.md call:
pwshFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires git, PowerShell, and gh CLI for PR context.
From compatibility in the SKILL.md frontmatter.
Evaluate PR Tests loads about 2.9k tokens when it runs. Until then it costs about 107 tokens; SKILL.md has 1,055 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from dotnet/maui at commit 7d38fd0, republished under its MIT licence (© dotnet). 1,055 words, ~2,949 tokens.
.claude/skills/evaluate-pr-tests/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Evaluates the quality, coverage, and appropriateness of tests added in a PR. Produces a structured report with actionable findings.
# Auto-detect PR and base branch
pwsh .github/skills/evaluate-pr-tests/scripts/Gather-TestContext.ps1
# With explicit base branch
pwsh .github/skills/evaluate-pr-tests/scripts/Gather-TestContext.ps1 -BaseBranch "origin/main"Run the script to get file categorization, convention checks, and anti-pattern detection:
pwsh .github/skills/evaluate-pr-tests/scripts/Gather-TestContext.ps1This produces a report at CustomAgentLogsTmp/TestEvaluation/context.md with:
Read the fix files to understand:
Read each test file and evaluate against all criteria below. For each criterion, provide a verdict (✅ Pass, ⚠️ Concern, ❌ Fail) with explanation.
Output a structured evaluation report (see Output Format below).
Question: Does the test exercise the actual code paths changed by the fix?
How to check:
Red flags:
Example — Good:
// Fix: CollectionView.SelectedItem setter now clears selection when set to null
// Test: Sets SelectedItem to null and verifies selection is cleared
App.Tap("SelectItem");
App.Tap("ClearSelection"); // Sets SelectedItem = null
var text = App.FindElement("SelectionStatus").GetText();
Assert.That(text, Is.EqualTo("None")); // Directly tests the fixExample — Bad:
// Fix: CollectionView.SelectedItem setter
// Test: Just checks CollectionView renders (doesn't test selection clearing)
App.WaitForElement("MyCollectionView");
Assert.That(true); // Proves nothing about the fixQuestion: Does the test cover boundary conditions, or only the happy path?
Check for these common gaps:
| Gap Type | What to Look For |
|---|---|
| Null/empty | Does the fix handle null? Is it tested? |
| Boundary values | Min, max, zero, negative, very large |
| Repeated actions | Does calling the action twice cause issues? |
| Platform-specific | Does the bug only occur on certain platforms? |
| Async/timing | Does the fix involve async code? Race conditions? |
| State transitions | Does the test cover before→after state changes? |
| Error paths | What happens when the operation fails? |
| Combination effects | Does the fix interact with other properties/features? |
How to suggest missing edge cases:
if (x == null), if (x <= 0), try/catch blocksQuestion: Is this the lightest test type that can verify the fix?
Preference order (lightest → heaviest):
| Priority | Type | When Appropriate | Project |
|---|---|---|---|
| ⭐ 1st | Unit Test | Pure logic, property changes, data transformations, binding behavior, event wiring | *.UnitTests.csproj |
| ⭐ 1st | XAML Test | XAML parsing, XamlC compilation, source generation, markup extensions | Controls.Xaml.UnitTests |
| ⭐⭐ 2nd | Device Test | Platform-specific rendering, native API interaction, handler mapping | *.DeviceTests.csproj |
| ⭐⭐⭐ 3rd | UI Test | User interaction flows, visual layout, screenshot comparison, end-to-end scenarios | TestCases.Shared.Tests |
Decision tree:
Does the test need to interact with visual UI elements?
YES → Is it checking visual layout/appearance?
YES → UI test (VerifyScreenshot) ✅
NO → Could the interaction be tested via handler/control API?
YES → Device test ⭐⭐
NO → UI test ✅
NO → Does it need a platform/native context?
YES → Device test ⭐⭐
NO → Does it test XAML parsing/compilation?
YES → XAML test ⭐
NO → Unit test ⭐Common "could be lighter" patterns:
| Current Test Does | Could Be Instead | Why |
|---|---|---|
| UI test: sets property, checks label text | Unit test | Property logic doesn't need UI |
| UI test: verifies event fires | Unit test | Event wiring is testable in isolation |
| UI test: checks control doesn't crash | Device test | Don't need Appium for crash testing |
| UI test: validates XAML binding | XAML test | Binding resolution is compile-time |
| Device test: checks property default | Unit test | Defaults don't need platform context |
Automated by the script. Review the script output for:
UI Tests:
IssueXXXXX.cs[Issue()] attribute on HostApp page[Category()] attribute — exactly ONE per test class (on the class or method, not both)_IssuesUITest base classWaitForElement before interactionsTask.Delay/Thread.Sleep#if ANDROID/#if IOSApplication.MainPage, Frame, Device.BeginInvokeOnMainThread)UITestEntry/UITestEditor for screenshot testsUnit Tests:
[Fact] or [Theory] attributes (xUnit)XAML Tests:
[Test] with [Values] XamlInflator parameterMauiXXXXXQuestion: Is this test likely to be flaky in CI?
| Risk Factor | Detection | Mitigation |
|---|---|---|
| Arbitrary delays | Task.Delay, Thread.Sleep | Use WaitForElement, retryTimeout |
| Missing waits | App.Tap without prior WaitForElement | Add explicit waits |
| Screenshot timing | VerifyScreenshot() without retryTimeout | Add retryTimeout: TimeSpan.FromSeconds(2) |
| Cursor blink | Entry/Editor in screenshot test | Use UITestEntry/UITestEditor |
| External URLs | WebView loading remote content | Use mock URLs or local content |
| Animation timing | Visual check after animation | Use retryTimeout |
| Global state | Test modifies Application.Current | Ensure cleanup in teardown |
Question: Does a similar test already exist?
Check the "Existing Similar Tests" section of the script output. If similar tests exist:
Question: Does the test run on all platforms affected by the fix?
Check the "Platform Scope Analysis" from the script:
.ios.cs compiles for both)Question: Are the assertions specific enough to catch regressions?
| Assertion Quality | Example | Verdict |
|---|---|---|
| ✅ Specific | Assert.That(label.Text, Is.EqualTo("Expected Value")) | Catches regression |
| ⚠️ Vague | Assert.That(label.Text, Is.Not.Null) | Too permissive |
| ❌ Meaningless | Assert.That(true) or no assertion | Proves nothing |
| ✅ Positional | Assert.That(rect.Y, Is.GreaterThan(safeAreaTop)) | Specific to layout fix |
| ⚠️ Brittle | Assert.That(rect.Y, Is.EqualTo(47)) | Magic number, will break |
Question: Do the files changed by the fix align with what the test exercises?
Red flags:
Issue12345 for a fix in CollectionView but only exercises Label renderingShell.cs but test only navigates a ContentPageProduce the evaluation report in this format:
## PR Test Evaluation Report
**PR:** #XXXXX — [Title]
**Test files evaluated:** [count]
**Fix files:** [count]
---
### Overall Verdict
[One of: ✅ Tests are adequate | ⚠️ Tests need improvement | ❌ Tests are insufficient]
[1-2 sentence summary of the most important finding]
---
### 1. Fix Coverage — [✅/⚠️/❌]
[Does the test exercise the code paths changed by the fix?]
### 2. Edge Cases & Gaps — [✅/⚠️/❌]
**Covered:**
- [edge case 1]
- [edge case 2]
**Missing:**
- [gap 1 — describe what should be tested and why]
- [gap 2]
### 3. Test Type Appropriateness — [✅/⚠️/❌]
**Current:** [UI Test / Device Test / Unit Test / XAML Test]
**Recommendation:** [Same / Could be lighter — explain why]
### 4. Convention Compliance — [✅/⚠️/❌]
[Summary from automated checks — list only issues found]
### 5. Flakiness Risk — [✅ Low / ⚠️ Medium / ❌ High]
[Specific risk factors identified]
### 6. Duplicate Coverage — [✅ No duplicates / ⚠️ Potential overlap]
[Similar existing tests found, if any]
### 7. Platform Scope — [✅/⚠️/❌]
[Does test coverage match the platforms affected by the fix?]
### 8. Assertion Quality — [✅/⚠️/❌]
[Are assertions specific enough to catch the actual bug?]
### 9. Fix-Test Alignment — [✅/⚠️/❌]
[Do the test and fix target the same code paths?]
---
### Recommendations
1. [Most important actionable recommendation]
2. [Second recommendation]
3. [...]| File | Description |
|---|---|
CustomAgentLogsTmp/TestEvaluation/context.md | Automated context report from script |
| Problem | Cause | Solution |
|---|---|---|
| No changed files detected | Wrong base branch | Use -BaseBranch explicitly |
| No fix files detected | All changes are tests | Expected for test-only PRs |
| AutomationId mismatch | HostApp and test out of sync | Update one to match the other |
| Convention check false positive | Script regex too broad | Ignore and note in report |
© dotnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts) in .github/skills/evaluate-pr-tests of dotnet/maui.
Open the folder on GitHubat commit 7d38fd0
Evaluate PR Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluate PR Tests this skilldotnet/maui | 23k | — | ~2.9k | Automated safety check: Pass | MIT | |
| Differential Security Reviewtrailofbits/skills | 7.4k | — | ~1.8k | Automated safety check: Notes | CC-BY-SA-4.0 | |
| Code Review WorkflowResgrid/Core | 229 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Reviewatelier-fashion/adlc-toolkit | 171 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Code Reviewpolyipseity/obsidian-terminal | 948 | — | ~1.6k | Automated safety check: Pass | AGPL-3.0 | |
| Core Components Code Reviewcore-ds/core-components | 137 | — | ~5.4k | Automated safety check: Pass | MIT |
trailofbits/skills
Reviews a pull request, commit or diff for security problems, using git history, caller counts and test coverage, and writes a markdown report.
Resgrid/Core
Structured code review workflow for .NET projects using Roslyn MCP tools.
atelier-fashion/adlc-toolkit
Multi-agent code review covering correctness, quality, architecture, test coverage, and security
polyipseity/obsidian-terminal
A skill your agent uses when reviewing PRs, code changes, or conducting code audits in obsidian-terminal.
core-ds/core-components
Review a Pull Request or diff in the @alfalab/core-components UI library — correctness bugs, public API/breaking changes, accessibility, keyboard/focus/pointer interaction, component states…
QwenLM/qwen-code
Runs a sandboxed, evidence-based check of one qwen-code pull request, proving its main change against the base build and writing a report with a machine-readable verdict.
dotnet/maui
Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.
dotnet/maui
Produces evidence-backed ship-readiness verdicts for .NET MAUI Servicing Releases and Previews, and drafts public-safe release handoff pages from the result.
dotnet/maui
Interprets pinned managed benchmark evidence for a dotnet/maui pull request and writes a narrative for the performance review workflow, without running or publishing anything.
dotnet/maui
Checks that a pull request's title and description match its implementation and reviews the code for best practices before merge, without posting anything.
dotnet/maui
Adds MAUI-specific guardrails on top of the maestro-cli skill and Maestro MCP tools for darc, BAR, and channel or feed lookups in dotnet/maui.
dotnet/maui
Adds dotnet/maui-specific context for investigating failing PR checks and broken nightly builds: pipelines, Helix logs, binlogs and merge-readiness verdicts.
Works with
Categories
Reviews the tests added in a pull request for fix coverage, quality, edge cases and test type, and recommends lighter test types where they would do. The skill reviews the tests added in a pull request and produces a structured report with actionable findings.ps1`, which writes a context report covering file categorization (fix files versus test files), convention checks for naming, attributes and anti-patterns, AutomationId consistency between the host app and the tests, similar existing tests and platform scope.
Evaluate PR Tests fits situations like: reviewing whether the tests in a PR actually cover the fix; checking if a lighter test type could replace a UI test; assessing test quality before merging a pull request.
Run `npx skills add dotnet/maui --skill evaluate-pr-tests -a claude-code`. Or copy the skill folder (.github/skills/evaluate-pr-tests in dotnet/maui) into .claude/skills/evaluate-pr-tests in your project. Claude Code loads it when a task matches its description.
Run `npx skills add dotnet/maui --skill evaluate-pr-tests -a codex`. Or copy the skill folder (.github/skills/evaluate-pr-tests in dotnet/maui) into .agents/skills/evaluate-pr-tests in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dotnet/maui --skill evaluate-pr-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-pr-tests, .gemini/skills/evaluate-pr-tests, .github/skills/evaluate-pr-tests and .opencode/skills/evaluate-pr-tests in your project.
Going by SKILL.md and its folder, Evaluate PR Tests needs PowerShell for the scripts in its folder and the command-line tools its instructions call (pwsh). Our summary lists: git, PowerShell and the gh CLI for PR context. Compatibility (from SKILL.md): Requires git, PowerShell, and gh CLI for PR context..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Evaluate PR Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Evaluate PR Tests: Differential Security Review (trailofbits/skills, 7.4k stars), Code Review Workflow (Resgrid/Core, 229 stars), Review (atelier-fashion/adlc-toolkit, 171 stars) and Code Review (polyipseity/obsidian-terminal, 948 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
dotnet (a GitHub organization, an official publisher) maintains it in dotnet/maui, which has 23,322 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 7, 2026.
Source: dotnet/maui on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.