Official agent skill

Code Review

by microsoft in microsoft/testfx

Review pull requests for MSTest and Microsoft.Testing.Platform with TestFx-specific checks.

OfficialMITAuto-check passedDevelopment

Install Code Review

skills CLI
$ npx skills add microsoft/testfx --skill code-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/testfx code-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/testfx.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/code-review .claude/skills/code-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
code-review
GitHub stars
1k
Token cost
~3.7k tokens
SKILL.md length
1,846 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Review pull requests for MSTest and Microsoft.Testing.Platform with TestFx-specific checks.

  • Works in 6 steps: Identify the product area from the… → Trace each behavior change through its… → Compare the PR title and description… → …
  • Every GitHub Copilot code review in this repository
  • SKILL.md covers Review process, Review publication, PR scope and review depth and Core TestFx checks, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Code Review is an agent skill from microsoft/testfx, published by the product's own GitHub organization. Review pull requests for MSTest and Microsoft.Testing.Platform with TestFx-specific checks. Use for every GitHub Copilot code review in this repository, with extra scrutiny for tests, public APIs, analyzers, MSBuild, localization, and agentic workflows.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Code review, Pull requests and Internationalization. The repository describes itself as: This repository holds the source code of Microsoft.Testing.Platform (MTP), a lightweight alternative to VSTest, as well as MSTest adapter and framework. The licence is MIT.

When your agent uses it

  • Every GitHub Copilot code review in this repository
  • With extra scrutiny for tests
  • Agentic workflows

Example prompts

  • “/code-review”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Identify the product area from the changed paths
  2. Trace each behavior change through its callers, tests, target frameworks,
  3. Compare the PR title and description with the actual diff. Verify that the
  4. Check changed tests against the test-review guidance below.
  5. Report only actionable findings caused by the PR. Prefer a small number of
  6. If the change is correct, do not invent a finding merely to leave feedback.

What it can do on your machine

Read from SKILL.md and the folder at commit 44b9dcc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Code Review loads about 3.7k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 1,846 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/testfx at commit 44b9dcc, republished under its MIT licence (© microsoft). 1,846 words, ~3,674 tokens.

Download SKILL.mdSave it as .claude/skills/code-review/SKILL.md (or your agent's skills folder).
name
code-review
description
Review pull requests for MSTest and Microsoft.Testing.Platform with TestFx-specific checks. Use for every GitHub Copilot code review in this repository, with extra scrutiny for tests, public APIs, analyzers, MSBuild, localization, and agentic workflows.

TestFx Code Review

Review the pull request as a maintainer of MSTest and Microsoft.Testing.Platform. Read .github/copilot-instructions.md first and treat it as authoritative.

Focus on defects introduced by the pull request. Do not report pre-existing issues, speculative concerns without a concrete failure mode, or style preferences already enforced by automation. Read enough surrounding code and tests to understand the changed behavior before commenting.

Review process

  1. Identify the product area from the changed paths:
    • src/Platform and related tests: Microsoft.Testing.Platform and extensions.
    • src/TestFramework: the public MSTest framework API.
    • src/Adapter: MSTest adapters and platform services.
    • src/Analyzers: Roslyn analyzers and code fixes.
    • src/Package/MSTest.Sdk, eng, and MSBuild files: build and packaging.
  2. Trace each behavior change through its callers, tests, target frameworks, shipped packages, and linked source files.
  3. Compare the PR title and description with the actual diff. Verify that the stated motivation, behavior change, compatibility impact, and validation match what the code does.
  4. Check changed tests against the test-review guidance below.
  5. Report only actionable findings caused by the PR. Prefer a small number of high-confidence findings over broad summaries or praise.
  6. If the change is correct, do not invent a finding merely to leave feedback.

Review publication

  • Publish review feedback as one pull-request review per run. Stage actionable line findings as inline review comments, then bundle them with the final review submission.
  • Put PR-level findings, scope and description feedback, dependency assessments, specialist-review summaries, and overflow findings in the final review body. Do not post separate top-level PR comments for review content.
  • Include a compact Confidence at a glance section with one color-coded, collapsed block per applicable review scope. Keep the detailed dimensions as an internal coverage checklist; do not publish one status row per dimension.
  • Lead with measurable reviewer actions when action is needed. Keep detailed line-level evidence in inline comments rather than repeating it in the final review body.
  • Do not duplicate a finding already covered by another live review thread. Reference the existing thread from the review body when it remains the only actionable item.
  • For automatic specialist checks that are fully clean, prefer noop when the main review already covers that dimension. Explicit slash-command reviews may still publish an informational COMMENT review.

PR scope and review depth

Be constructively critical. Challenge the change's assumptions, failure modes, and claimed validation rather than accepting the implementation or description at face value. Criticism must still identify concrete evidence and impact; do not manufacture objections for the sake of sounding rigorous. Every criticism must meet the concrete-scenario and observable-consequence bar in Finding quality.

  • Check whether the PR combines independent features, refactors, fixes, or formatting changes that have different motivations or could be reviewed, reverted, and shipped separately. When this materially obscures behavior or increases review risk, recommend splitting the PR and identify the distinct change groups.
  • Check that every material diff is explained by the PR description and that every claimed behavior is implemented. Flag hidden scope, stale claims, understated compatibility or operational impact, and validation claims not supported by the changed tests or available evidence. A missing or brief description is not itself a defect; raise it only when it hides material behavior, compatibility impact, or an unsupported claim.
  • Distinguish necessary supporting changes from unrelated cleanup. Do not request a split merely because a coherent change spans many files or product layers.
  • Put scope, description-alignment, split, and additional-review feedback only in the top-level review summary, never on an arbitrary changed line.
  • Request an independent multi-model review sparingly, only for high-blast- radius changes such as:
    • Wire-level IPC contracts or serialization protocols shared with external clients.
    • Structural redesigns of core test-execution scheduling, synchronization, cancellation, or ExecutionContext propagation.
    • New process-launch, arbitrary-code-execution, privilege, or other security boundaries.
    • Cross-product changes with independent compatibility or rollback risks.
  • Do not request a multi-model review for routine public API additions, bug fixes, analyzer changes, adapter features, or test-only changes. Put the request in the review summary and name the exact failure mode, compatibility concern, race, or attack surface the additional models should examine.

Core TestFx checks

  • Preserve backward compatibility for public APIs, protocols, command-line options, package behavior, and analyzer diagnostics.
  • Treat overload additions as potentially source-breaking. Check representative call shapes and older supported language versions, not only the repository's preview language version.
  • New public API must be minimal, documented, MUST NOT use init, and be recorded in the relevant PublicAPI.Unshipped.txt. Also check internal API baselines in projects that track them.
  • Every API marked [Experimental] must include this sentence in its XML documentation <remarks>: This API is experimental. It may change, break, or be removed at any time without notice.
  • Check every target framework affected by the change. Do not assume an API available on modern .NET exists on older targets.
  • Review concurrency and lifecycle ordering carefully. Test execution is parallel, and shared mutable state, cancellation, disposal, and ExecutionContext flow are common correctness boundaries.
  • For shared or linked source, identify every consuming project and ensure API baselines and behavior remain consistent in all of them.
  • User-facing strings belong in resources. Never accept manually edited XLF files. Every {Locked="..."} token must occur verbatim in the corresponding value; verify it does not accidentally lock a substring of a translatable word.
  • CLI option changes must update the matching --help and --info acceptance expectations, including punctuation and alphabetical ordering.
  • Packable projects need valid package descriptions and PACKAGE.md.
  • Never accept manual changes under eng/common/; those files are mirrored from dotnet/arcade and overwritten by automation.
  • Do not weaken ApplicationStateGuard.Unreachable(), Debug.Assert, or explicit invariant throws into silent fallbacks, warnings, or empty returns without a concrete external trigger or failing test that disproves the invariant.
  • Workflow, script, issue-template, and policy changes that create or triage issues must use native GitHub Issue Types and must not introduce the deprecated type/bug, type/feature, or type/task labels.
  • Do not allow untracked TODO comments.
Show full SKILL.md (884 more words)Show less

Security checks

Review changed trust boundaries for concrete security regressions. Pay particular attention when the change handles paths, process arguments, environment variables, protocol messages, serialized data, logs, artifacts, credentials, reflection, extensions, or arbitrary user test code.

  • Ensure untrusted file and directory names cannot escape the intended root through traversal, rooted paths, alternate separators, or symlink behavior.
  • Ensure process launches preserve argument boundaries and do not concatenate untrusted values into a command line, shell command, or executable path. Check executable resolution, working-directory selection, inherited environment variables, and assembly or plugin search paths.
  • Validate untrusted protocol and serialized input before using it, including message type, required fields, payload and collection bounds, and unsupported values. Reject unsafe polymorphic or binary deserialization of untrusted data.
  • Treat reflection, user callbacks, test methods, data sources, extensions, and assembly discovery as hostile-code boundaries. Ensure exceptions are contained appropriately, cancellation and timeouts remain enforceable, and user-controlled input cannot cause unbounded memory or resource growth.
  • Do not expose secrets, tokens, environment values, private paths, or sensitive test data through logs, exceptions, reports, artifacts, telemetry, or temporary files.
  • Check that temporary files and artifacts use safe locations, appropriate access, collision-resistant names, and reliable cleanup.
  • Treat workflow permission, network access, action pin, and secret-handling changes as security boundaries; reject unnecessary privilege expansion or mutable third-party references.
  • For dependency changes, check whether the new package or version introduces a known vulnerability, unexpected runtime asset, or broader transitive surface. Use authoritative evidence such as official advisories, NuGet metadata, and upstream release notes; preserve uncertainty rather than guessing.

Report a security finding only when the changed code creates a plausible attack or disclosure scenario with an observable consequence. Defensive coding belongs at actual external boundaries; do not request blanket hardening of trusted internal invariants without a concrete external trigger.

For a suspected vulnerability or security-driven dependency update, do not put proofs of concept, attacker payloads, exploit steps, CVE details, affected version ranges, or other actionable vulnerability details in public review comments or PR metadata. Leave at most a minimal public note and follow SECURITY.md and the repository's private reporting process.

Test changes

For changed test code, apply the relevant rules from:

  • .agents/skills/test-anti-patterns/SKILL.md
  • .agents/skills/assertion-quality/SKILL.md
  • .agents/skills/test-gap-analysis/SKILL.md

Use those files as review checklists, not as a request for a repository-wide audit. Keep analysis scoped to the changed behavior and its directly related tests.

In particular, verify:

  • Tests would fail if the production behavior were wrong; reject missing, tautological, self-comparing, or only-trivial assertions.
  • Async assertions and operations are awaited.
  • Tests are isolated under parallel execution and restore environment, culture, static state, files, and other process-wide state.
  • Synchronization is deterministic. Avoid sleeps and fragile duration assertions when a marker, event, or rendezvous can prove the behavior.
  • Acceptance tests that share a generated mutable asset are marked [DoNotParallelize]; do not require it when the state is execution-context local or otherwise isolated.
  • Duration output uses AcceptanceAssert.DurationPattern rather than assuming a millisecond-only format.
  • The test framework and assertion library match the project's BannedSymbols.txt and repository conventions.
  • MSTest framework unit tests use TestFramework.ForTestingMSTest; MTP and analyzer tests use MSTest; adapter tests use their established assertion library.
  • Test changes cover negative paths, boundaries, cleanup, cancellation, and failure behavior relevant to the production change.

Specialist review routing

Apply these additional repository resources when their paths are changed:

  • MSBuild files (*.props, *.targets, *.csproj, SDK and NuGet build extensions): use the rule catalog in .github/agents/msbuild-reviewer.agent.md. Ignore its operating modes, finding cap, orchestration instructions, and output contract; the GitHub code-review service owns review formatting and posting.
  • Broad TestFx architecture, runtime, API, performance, compatibility, security-boundary, IPC, process-launch, serialization, and artifact-handling changes: use the applicable review dimensions in .github/agents/expert-reviewer.agent.md. Ignore workflow-specific posting, attribution, and safe-output instructions in that file.
  • Agentic workflow Markdown: require strict compilation, regenerated lock files, unchanged trusted action pins, and the repository action-pin audit.
  • Analyzer changes:
    • Verify diagnostic IDs, severity, generated-code behavior, false-positive risk, analyzer release tracking, and code-fix registration and properties.
    • For every mapping from a source API or annotation domain to a target API or runtime domain, classify the mapping as exact, compatible/coarser, or unrepresentable. Check all source-value polarities and supported target versions; do not infer equivalence from one successful case.
    • Separate compile-time annotation semantics from runtime enforcement. Verify what the analyzer can prove from symbols and metadata independently from what the target framework or platform actually enforces at execution time.
    • Check descriptor wording against the classification. Use "equivalent" only for exact mappings; describe lossy but behaviorally acceptable mappings as compatible and make any semantic loss explicit.
    • Every diagnostic without a code fix must have a safe, concrete manual edit that clears the diagnostic while preserving the relevant behavior. If no such edit exists for a valid triggering program, question whether the diagnostic is actionable rather than treating the expected diagnostic as proof of correctness.

Finding quality

Every finding must:

  • Identify a concrete scenario that triggers the problem.
  • Explain the observable consequence.
  • Point to a changed line for code findings.
  • Put PR-level scope, description, split, and multi-model-review feedback in the overall review summary and identify the specific unsupported claim, unexplained change group, or high-risk boundary.
  • Recommend the smallest safe correction.
  • Use severity proportional to impact; do not elevate maintainability or style suggestions into correctness findings.

Do not submit an approval on behalf of repository maintainers. A clean review may recommend approval in its summary while remaining a comment review.

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/code-review of microsoft/testfx.

Open the folder on GitHubat commit 44b9dcc

Compare with similar skills

Code Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Code Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Code Review this skillmicrosoft/testfx1k—~3.7kAutomated safety check: PassMIT
Code Reviewmicrosoft/vscode-containers141—~415Automated safety check: PassCustom licence
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Understand Diff AnalysisEgonex-AI/Understand-Anything85k1 repos~1.4kAutomated safety check: PassMIT
WooCommerce Code Reviewwoocommerce/woocommerce11k3 repos~1.1kAutomated safety check: PassCustom licence
Open Code Review CLIalibaba/open-code-review44k—~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • Code Review

    microsoft/vscode-containers

    Official

    Review a pull request in this repository (the vscode-containers / Container Tools extension) when the automated reviewer has already supplied the diff and changed files.

    141 GitHub stars~415 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Understand Diff Analysis

    Egonex-AI/Understand-Anything

    Reads your git changes or a pull request against a prebuilt knowledge graph of the project to explain what changed, which components are affected and what is risky.

    85k GitHub starsUsed in 1 repo~1.4k tokens
    DevelopmentAuto-check passed
  • WooCommerce Code Review

    woocommerce/woocommerce

    Reviews WooCommerce code changes against the project's standards, flagging backend PHP architecture, naming, documentation, data integrity and testing violations.

    11k GitHub starsUsed in 3 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Open Code Review CLI

    alibaba/open-code-review

    Runs the ocr command-line tool to review Git changes, a commit or a branch comparison with an AI model, returning line-level comments and optionally applying fixes.

    44k GitHub stars~3.1k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Official

    Runs a loop on a GitHub pull request: fetch review state, triage comments into actions, implement them and resolve threads, repeating until nothing actionable is left.

    48k GitHub stars~2.2k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from microsoft/testfx

All 44 skills in this repo
  • Official

    Guide for organizing MSBuild infrastructure with Directory.Build.props, Directory.Build.targets, Directory.Packages.props, and Directory.Build.rsp.

    1k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Binlog Failure Analysis

    microsoft/testfx

    Official

    Analyze MSBuild binary logs to diagnose build failures. An agent skill from microsoft/testfx.

    1k GitHub starsUsed in 3 repos~730 tokens
    Auto-check passed
  • Coverage Analysis

    microsoft/testfx

    Official

    Project-wide code coverage and CRAP (Change Risk Anti-Patterns) score analysis for .NET projects.

    1k GitHub stars~7.3k tokensUpdated today
    Auto-check passed
  • Incremental Build

    microsoft/testfx

    Official

    Guide for optimizing MSBuild incremental builds. An agent skill from microsoft/testfx.

    1k GitHub starsUsed in 3 repos~3.7k tokens
    Auto-check passed
  • Msbuild Modernization

    microsoft/testfx

    Official

    Guide for modernizing and migrating MSBuild project files to SDK-style format.

    1k GitHub starsUsed in 3 repos~4.3k tokens
    Auto-check passed
  • Msbuild Antipatterns

    microsoft/testfx

    Official

    Catalog of MSBuild anti-patterns with detection rules and fix recipes.

    1k GitHub stars~4.5k tokensUpdated today
    Auto-check passed

Categories

Questions about Code Review

What does Code Review do?

Review pull requests for MSTest and Microsoft.Testing.Platform with TestFx-specific checks. Code Review is an agent skill from microsoft/testfx, published by the product's own GitHub organization.Platform with TestFx-specific checks.

When should I use Code Review?

Code Review fits situations like: every GitHub Copilot code review in this repository; with extra scrutiny for tests; agentic workflows.

How do I install Code Review in Claude Code?

Run `npx skills add microsoft/testfx --skill code-review -a claude-code`. Or copy the skill folder (.github/skills/code-review in microsoft/testfx) into .claude/skills/code-review in your project. Claude Code loads it when a task matches its description.

How do I install Code Review in Codex?

Run `npx skills add microsoft/testfx --skill code-review -a codex`. Or copy the skill folder (.github/skills/code-review in microsoft/testfx) into .agents/skills/code-review in your project. Codex loads it when a task matches its description.

Can I use Code Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/testfx --skill code-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-review, .gemini/skills/code-review, .github/skills/code-review and .opencode/skills/code-review in your project.

What does Code Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Code Review is instructions for the agent only.

Does Code Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Code Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Code Review use?

Code Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Code Review use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Code Review?

Skills that share tags, products or a category with Code Review: Code Review (microsoft/vscode-containers, 141 stars), PR Babysitter (openinterpreter/openinterpreter, 69k stars), Understand Diff Analysis (Egonex-AI/Understand-Anything, 85k stars) and WooCommerce Code Review (woocommerce/woocommerce, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Code Review?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/testfx, which has 1,047 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 7, 2026.

Source: microsoft/testfx on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.