Official agent skill

Grade Tests

by dotnet in dotnet/skills

Assess a curated list of tests and produce a PR-ready table with a primary Pass, Failed, Uncertain, or Not applicable result plus A-F quality detail for every resolved test; Uncertain and Not…

OfficialMITAuto-check passedMobile

Install Grade Tests

skills CLI
$ npx skills add dotnet/skills --skill grade-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dotnet/skills grade-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dotnet/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/dotnet-test/skills/grade-tests .claude/skills/grade-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
grade-tests
GitHub stars
5.6k
Used in
1 other repo
Token cost
~5.2k tokens
SKILL.md length
2,652 words
Files
1
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Assess a curated list of tests and produce a PR-ready table with a primary Pass, Failed, Uncertain, or Not applicable result plus A-F quality detail for every resolved test; Uncertain and Not…

  • Works in 7 steps: Validate the input → Detect language and load extension → Resolve the test bodies → …
  • Modified tests supplied as methods
  • SKILL.md covers Why a Decision Result Plus…, When to Use, When Not to Use and Inputs, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Grade Tests is an agent skill from dotnet/skills, published by the product's own GitHub organization. Assess a curated list of tests and produce a PR-ready table with a primary Pass, Failed, Uncertain, or Not applicable result plus A-F quality detail for every resolved test; Uncertain and Not applicable omit the grade. USE FOR new or modified tests supplied as methods, bodies, file spans, or a bounded PR diff. Polyglot: .NET, Python, TS/JS, Java, Go, Ruby, Rust, Swift, Kotlin, PowerShell, C++. DO NOT USE FOR: suite-wide audits (use test-quality-auditor or test-anti-patterns), writing or fixing tests, or measuring…

Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Mobile, covering Android development and iOS development. It works with .NET, Python, C++ and Java. The repository describes itself as: Repository for skills to assist AI coding agents with .NET and C. The licence is MIT.

When your agent uses it

  • Modified tests supplied as methods
  • A bounded PR diff
  • : suite-wide audits (use test-quality-auditor
  • Test-anti-patterns)

Example prompts

  • “/grade-tests”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Validate the input
  2. Detect language and load extension
  3. Resolve the test bodies
  4. Score each resolved test
  5. Assign the decision result
  6. Build the note
  7. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 8d670fa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Grade Tests loads about 5.2k tokens when it runs. Until then it costs about 135 tokens; SKILL.md has 2,652 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~135
When it runs · the whole SKILL.md, loaded when a task matches
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dotnet/skills at commit 8d670fa, republished under its MIT licence (© dotnet). 2,652 words, ~5,190 tokens.

Download SKILL.mdSave it as .claude/skills/grade-tests/SKILL.md (or your agent's skills folder).
name
grade-tests
description
Assess a curated list of tests and produce a PR-ready table with a primary Pass, Failed, Uncertain, or Not applicable result plus A-F quality detail for every resolved test; Uncertain and Not applicable omit the grade. USE FOR new or modified tests supplied as methods, bodies, file spans, or a bounded PR diff. Polyglot: .NET, Python, TS/JS, Java, Go, Ruby, Rust, Swift, Kotlin, PowerShell, C++. DO NOT USE FOR: suite-wide audits (use test-quality-auditor or test-anti-patterns), writing or fixing tests, or measuring coverage.
license
MIT

Grade Tests

Assess a curated list of test methods and produce a compact, PR-comment-friendly report. The primary result is one of Pass, Failed, Uncertain, or Not applicable; an A-F quality grade remains secondary diagnostic information. The skill does not discover tests on its own — the caller (typically a PR automation workflow or a human reviewer holding a specific list) provides the tests or a bounded diff to assess.

Language-specific guidance: Call the test-analysis-extensions skill to discover available extension files, then read the file matching the target codebase's language and framework (e.g., extensions/dotnet.md, extensions/python.md, extensions/typescript.md, extensions/go.md). You MUST read the relevant extension file before scoring assertions or anti-patterns, because assertion APIs and idiomatic patterns differ significantly across frameworks.

Why a Decision Result Plus Quality Detail

PR reviewers need a simple answer to does this test need follow-up? The four-state result provides that decision; the existing A-F rubric explains its quality and severity.

When to Use

  • A PR automation workflow needs to post a decision on the tests introduced or modified in a pull request.
  • A reviewer has a specific list of tests (a file, a class, a method list, or a diff hunk) and wants per-test follow-up decisions rather than a suite report.
  • A maintainer wants to triage which of N tests in a contribution deserve follow-up improvements, with quality grades for resolved tests.

When Not to Use

  • The caller wants a full suite audit or comparative metrics — use test-anti-patterns (pragmatic) or test-smell-detection (formal) and let the test-quality-auditor agent orchestrate.
  • The caller wants to write new tests — use code-testing-generator (any language) or writing-mstest-tests (MSTest specifically).
  • The caller wants to measure code coverage or CRAP scores — use coverage-analysis or crap-score (.NET only).
  • The caller wants to fix issues directly in test code — invoke the appropriate editing skill.
  • No specific list of tests is provided. Do not try to grade every test in the workspace; ask the caller for an explicit list or scope.

Inputs

InputRequiredDescription
Test methodsYesA scope to grade. Provide one of: (a) an explicit list of test method names (fully-qualified, e.g. Namespace.ClassName.TestMethodName); (b) one or more file paths plus an explicit instruction to grade every test declared in those files; or (c) a diff hunk / PR identifier whose changed tests should be graded. File paths are recommended but optional when method names are unambiguous in the workspace. Ambiguous requests like "grade my tests" with no scope are rejected up-front (see Step 0); this skill is for curated input and does not auto-grade an entire workspace.
Test bodies / spansRecommendedThe exact source lines for each test method. If omitted, read them from the listed files.
Production codeNoThe code under test, for judging whether assertions cover the meaningful behaviors. When unavailable, mark relevant findings as "Unverified" rather than guessing.
Diff contextNoWhen grading PR changes, the unified diff for each test method helps focus on what actually changed.
Step 0: Validate the input

Before doing anything else, check that the caller provided one of:

  1. An explicit list of test method names, or
  2. One or more file paths plus an explicit instruction to grade every test declared in those files (e.g., "grade every test in OrderTests.cs"), or
  3. A diff hunk or PR identifier whose changed tests should be graded.

If the request is ambiguous (e.g., "Grade my tests", "Are these tests any good?" with no scope, "Review the test suite"), do not load extensions, do not read files, and do not grade anything. Reply with a short message asking the caller to provide an explicit list / file(s) / diff, and optionally point them at test-quality-auditor agent or test-anti-patterns skill for full-suite analysis. Stop there.

If a valid bounded scope resolves to zero eligible tests, return Not applicable with a short explanation and no invented rows.

Workflow

Step 1: Detect language and load extension

Identify the target codebase's language and test framework from the file extensions and the test method markers in the provided list. Call the test-analysis-extensions skill and read the matching extension file (e.g., extensions/dotnet.md for MSTest/xUnit/NUnit/TUnit, extensions/python.md for pytest, extensions/typescript.md for Jest/Vitest, extensions/go.md for the standard testing package). If the input contains tests from multiple languages, load each relevant extension and grade each test using its language's conventions.

Step 2: Resolve the test bodies

For each entry in the input list:

  1. If the test body is provided inline, use it directly.
  2. Otherwise read the file at the given path and locate the method by its fully-qualified name. Capture the full method body, including attributes / decorators / fixtures and any helper code that the test calls.
  3. If a requested method cannot be found, record it as Uncertain — method not found with no quality grade and continue. Never invent a body to grade. A missing requested method requires human review; it is not the same as a valid scope containing no tests.
Step 3: Score each resolved test

Start every test at grade A (score band 90–100), then apply deductions strictly for observable issues in the captured body. Do not deduct for hypothetical concerns (e.g., "could have more negative assertions") unless the production code clearly demands them and the production code is available.

When production code is unavailable, grade observable issues in the test body normally, but do not infer missing behaviors or deduct for them. State Production-dependent behavior coverage: Unverified once in the summary so the reader can distinguish test-body findings from claims that require source code.

Three sub-dimensions

Compute three sub-grades (each A–F) that together drive the overall grade.

A. Assertion strength

Read the loaded language extension's assertion API list and classify every assertion in the test body. Score from highest to lowest:

Sub-gradePattern
AAt least one meaningful value assertion (equality / structural / exception / state) plus, where appropriate, additional checks (negative, type, collection contents). Mock-call verifications (Verify, toHaveBeenCalledWith, Should -Invoke) and bare assertion forms (pytest assert, Go if got != want { t.Errorf(...) }, Rust assert!()) count as real assertions.
BOne clear meaningful assertion that verifies the behavior under test.
COnly trivial assertions (single IsNotNull / toBeDefined / assert x is not None), or assertions that check a single field while the operation produces a richer result.
DOne self-referential / tautological assertion (Assert.AreEqual(x, x), assert dto.name == dto.name, round-trip identity without a non-trivial input), or broad exception assertions (Assert.ThrowsException<Exception>).
FNo assertions at all; all assertions are always-true literals (Assert.IsTrue(true), assert True, expect(true).toBe(true)) — these verify nothing and are equivalent to having no assertions; or all assertions are silently un-awaited (e.g., expect(promise).resolves.toBe(x) without await/return, async TUnit/xUnit Assert.ThrowsAsync without await, pytest-asyncio with un-awaited coroutine).

Exception and error-path tests (Assert.ThrowsException<T>, constrained pytest.raises, expect(fn).toThrow, assertThrows, #[should_panic], Should -Throw, EXPECT_THROW, or Go code that verifies an expected non-nil error) are complete on their own. Give Assertion strength A when the test checks the exact promised error condition for its stated scope. Do not deduct for having only that assertion, and do not require an error-message assertion unless the message is part of the documented contract. A Go happy-path test that only checks err == nil while discarding a meaningful returned value is still C because it does not verify the successful result.

B. Structure & focus
Sub-gradePattern
AClear Arrange-Act-Assert (or Given-When-Then) separation. Single behavior under test. Body under ~30 lines. Setup uses framework conventions.
BOne mild structural issue (slightly long body, missing blank lines between phases) but intent is clear.
CMultiple behaviors mixed in one test, or AAA phases interleaved enough to slow comprehension.
DConditional logic in the test (if/switch driving assertions) — except for idiomatic Go/Rust table-driven sub-test loops; or test relies on previous test state (ordering dependency).
FTest exceeds ~60 lines and verifies multiple unrelated behaviors; or shares mutable state with other tests through statics/globals without reset.
C. Anti-pattern hygiene

Scan against the catalog below. The Anti-pattern sub-grade is computed in two passes and combined deterministically:

  1. Hard ceiling pass. Every Critical or High finding sets a maximum sub-grade (F, D, or C as labeled). Take the worst ceiling across all matched Critical/High findings — these do not accumulate (a single F finding caps the sub-grade at F regardless of how many other Critical/High findings are present).
  2. Medium-deduction pass. Start from A, then for each Medium finding deduct one sub-grade level (A→B, B→C, C→D, D→F). These do accumulate across findings.

The final Anti-pattern sub-grade is the worse of the two passes (i.e., min(hard_ceiling, A − medium_count)). Low findings never affect the grade — mention them in the note only.

Examples (Critical/High and Medium counts → Anti-pattern sub-grade):

  • Zero Critical/High, 1 Medium → B (A − 1)
  • Zero Critical/High, 3 Medium → D (A − 3)
  • One C-ceiling (e.g., over-mocking), 0 Medium → C
  • One C-ceiling, 2 Medium → C (min(C, A − 2 = C) = C; a third Medium tips to D)
  • One F-finding (e.g., swallowed exception) plus any number of Medium → F

Critical (drop straight to F or D)

  • No assertions at all → F (also drives Assertion sub-grade to F)
  • Swallowed exceptions: try { … } catch { } (.NET), bare except: pass (Python), try { … } catch (e) {} (JS/TS/Java), defer recover() without re-panic (Go), rescue StandardError with no assertion (Ruby), empty catch (Kotlin/Swift) → F
  • Assert-in-catch pattern (Assert.Fail(ex.Message) instead of Assert.ThrowsException) → D
  • Always-true literal assertions (Assert.IsTrue(true), assert True, expect(true).toBe(true)) → F (verifies nothing; also drives Assertion sub-grade to F)
  • Self-referential / tautological assertions on bound values (Assert.AreEqual(x, x), assert dto.name == dto.name) → D
  • Commented-out assertions → D

High (drop one or two sub-grades)

  • Wall-clock sleep used for synchronization: Thread.Sleep, Task.Delay, time.sleep, setTimeout-based wait, Thread.sleep, time.Sleep, sleep, std::thread::sleep, Start-Sleep, std::this_thread::sleep_for (in a unit test) → D
  • Unseeded randomness, wall-clock reads without abstraction (DateTime.Now, datetime.now(), Date.now(), System.currentTimeMillis(), time.Now(), Time.now, Instant::now(), Get-Date, system_clock::now) → D
  • Hard-coded environment-dependent paths (C:\…, /tmp/…, network hosts) → D
  • Ordering dependency on mutable static / package globals → D
  • Broad exception assertion (Assert.ThrowsException<Exception>, pytest.raises(Exception), expect(fn).toThrow(Error) without matcher, #[should_panic] without expected = "…", Should -Throw without -ExpectedMessage, EXPECT_ANY_THROW) → C
  • Over-mocking: more mock setup lines than test logic, or verifying exact call sequences instead of outcomes → C
  • Implementation coupling: reflection on private members, casting to internal types to access state → C

Medium (drop one sub-grade)

  • Poor name: Test1, TestMethod, test, single-word name that says nothing about scenario or expected outcome (judge against the language extension's convention) → drop one sub-grade
  • Magic values: unexplained 42, "foo", 0x1234 in arrange/assert without naming or comment → drop one sub-grade
  • Giant test (>30 lines covering a single behavior) → drop one sub-grade
  • Assertion messages that just repeat the assertion text → drop one sub-grade
  • Missing AAA / GWT separation when the test is non-trivial → drop one sub-grade

Low (note only, no deduction)

  • Unused setup/teardown hooks; print debugging left in (Console.WriteLine, print, console.log, System.out.println, fmt.Println, puts, dbg!, Write-Host, std::cout); inconsistent naming versus siblings; leftover TODO comments. Mention in the note column but do not deduct.
Show full SKILL.md (887 more words)Show less
Combining sub-grades

Convert sub-grades to numeric points: A=4, B=3, C=2, D=1, F=0.

  • Overall score band = weighted average: 0.45 × Assertion + 0.30 × Anti-pattern + 0.25 × Structure
  • Map to letter:
    • ≥ 3.5 → A (band 90–100)
    • ≥ 2.8 → B (band 80–89)
    • ≥ 2.0 → C (band 70–79)
    • ≥ 1.2 → D (band 60–69)
    • < 1.2 → F (band 0–59)
  • The overall grade is capped at the worst sub-grade — if any sub-grade is F, the overall grade is F; if the worst sub-grade is D, the overall grade is at most D; and so on. A test that fails on any one dimension cannot earn a higher overall grade than that dimension.

Report the letter grade and the score band (not a single 0–100 number). False precision invites bikeshedding; bands keep the conversation focused on the rubric.

Step 4: Assign the decision result

The grade summarizes strength; the result says whether follow-up exists. An actionable improvement is an evidence-backed change to the test, setup, or fixtures. Assign exactly one:

  • Pass — no actionable improvement; positive/context-only notes are allowed.
  • Failed — at least one actionable improvement, regardless of grade.
  • Uncertain — missing evidence prevents a decision and needs human review.
  • Not applicable — a valid scope contains no eligible tests; normally an overall result with no rows.

Do not derive status from grade: a complete focused test can be B / Pass, while debug output can make an otherwise excellent test A / Failed. Use Uncertain for an unresolved body, unsupported construct, or essential missing contract—not merely absent production code. A definite finding wins over uncertainty.

Step 5: Build the note

Use one sentence (target ≤ 120 characters) for the most important reason: No issues found., Only checks IsNotNull; add value verification., or Method body could not be resolved; human review is required. Do not invent a weakness to justify a grade or Failed result.

Step 6: Report

Produce two sections.

1. Summary

Begin with **Result: <Pass|Failed|Uncertain|Not applicable>**, then give result counts and the highest-priority action. Aggregate using Failed → Uncertain → Pass → Not applicable. For Not applicable, explain the empty scope and omit the table.

2. Per-test table
markdown
| Test | Result | Quality | Notes |
|------|--------|---------|-------|
| `Namespace.ClassName.Test_Method_Condition_Expected` | Pass | A (90–100) | No issues found. |
| `Namespace.ClassName.Test_Other` | Failed | C (70–79) | Only `IsNotNull`; add value verification. |
| `Namespace.ClassName.Test_Missing` | Uncertain | — | Method body could not be resolved; human review is required. |

Caps and ordering:

  • If the table would exceed 50 rows, show Failed tests first, then Uncertain tests, then a sample of Pass tests. Wrap overflow in a collapsed <details> block.
  • Within the same result, order by quality from worst to best, then by file path and method name for determinism.
  • If the diff context is provided, prefix each test name with a (new) or (modified) marker.

If multiple languages are present, produce one table per language and prefix each section with the language name and framework.

Validation

  • Every test in the input list appears in the table (or is recorded as Uncertain — method not found).
  • Every resolved test has Pass or Failed plus A-F quality detail.
  • Uncertain is an evidence gap; Not applicable is a valid empty scope.
  • Every grade is justified by at least one observable signal in the captured body — no speculative deductions.
  • Trivial-assertion tests are flagged only when the only assertion is trivial (a null check before a meaningful assertion is not trivial).
  • Exception-only tests are not penalized for low assertion count.
  • Mock-call verifications and bare assertion forms count as real assertions of the appropriate category.
  • Boolean assertions on meaningful properties (Assert.IsTrue(result.IsValid)) are not classified as always-true; only literal true/false constants are.
  • Self-referential assertions are flagged separately from normal equality assertions.
  • Idiomatic patterns are not flagged: Go/Rust table-driven sub-tests, pytest bare assert, Go if got != want { t.Errorf(...) }, JS/TS expect(mock).toHaveBeenCalledWith(...).
  • Async test pitfalls (un-awaited resolves/rejects/ThrowsAsync, pytest-asyncio without await) drop the Assertion sub-grade to F.
  • The summary leads with the highest-leverage observation, not a recap of the table.

Common Pitfalls

PitfallSolution
Grading every test in the workspace when no list is providedAsk the caller for the explicit list; this skill is for curated input.
Inflating deductions to justify the gradeStart at A; deduct only for observable issues.
Penalizing exception tests for low assertion countException assertions are complete on their own.
Downgrading a focused Go error-path test because it checks only err != nilExpected-error existence is the observable contract for that scope; keep it at A unless the production contract requires a specific error identity or message.
Treating IsNotNull before a value assertion as trivialOnly flag when the null check is the only assertion.
Treating any Boolean assertion as effectively assertion-freeOnly always-true literals (Assert.IsTrue(true), assert True) are; meaningful Assert.IsTrue(result.IsValid) is a real assertion.
Flagging Go/Rust table-driven loops as conditional logicThey are idiomatic; do not deduct.
Treating pytest bare assert or Go if got != want { t.Error… } as missing-frameworkBoth are canonical; count in the correct assertion category.
Penalizing tests when production code is unavailableMark concerns about uncovered behaviors as Unverified and do not deduct.
Using a fake-precise score (e.g., 87/100)Use the score band only — 90–100, 80–89, 70–79, 60–69, 0–59.
Spilling a 500-row table into a PR commentApply the row cap from Step 6; collapse extras into <details>.
Re-reporting an existing finding three times under different categoriesPick the most fitting category and report once.
Inventing weaknesses for A-grade tests to make the note "balanced"If a test is clean, the note may simply read No issues found.
Mapping status from grade or commentsFail only for actionable improvements; a B can Pass and an A can Fail.
Confusing Uncertain and Not applicableEvidence gaps are Uncertain; a valid empty scope is Not applicable.

© dotnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/dotnet-test/skills/grade-tests of dotnet/skills.

Open the folder on GitHubat commit 8d670fa

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in dotnet/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Grade Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Grade Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Grade Tests this skilldotnet/skills5.6k1 repos~5.2kAutomated safety check: PassMIT
Code Revieweralirezarezvani/claude-skills28k1 repos~1.6kAutomated safety check: PassMIT
Build Teaql Appteaql/teaql-agent-kit2.8k—~4.6kAutomated safety check: PassMIT
Fory Version Bumpapache/fory4.6k—~1.1kAutomated safety check: PassApache-2.0
Fory Performance Optimizationapache/fory4.6k—~2.2kAutomated safety check: PassApache-2.0
Code Testing Extensionsmicrosoft/testfx1k2 repos~930Automated safety check: PassMIT

Similar skills

  • Code Reviewer

    alirezarezvani/claude-skills

    Code review automation for TypeScript, JavaScript, Python, Go, Swift, Kotlin, C, .NET, Java, C, C++, Rust, Ruby, PHP, and Dart/Flutter.

    28k GitHub starsUsed in 1 repo~1.6k tokens
    DevelopmentAuto-check passed
  • Build Teaql App

    teaql/teaql-agent-kit

    Build or change a TeaQL application in Java, Rust, Go, Swift, Python, C/.NET, or TypeScript, including Kotlin/JVM applications that consume Java-generated libraries.

    2.8k GitHub stars~4.6k tokensUpdated 10 days ago
    MobileAuto-check passed
  • Bump Apache Fory release or post-release development versions across Java, Kotlin, Scala, Python, Rust, Go, C++, C, Dart, JavaScript, Swift, integration tests, examples, and source docs.

    4.6k GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed
  • Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala).

    4.6k GitHub stars~2.2k tokensUpdated today
    MobileAuto-check passed
  • Code Testing Extensions

    microsoft/testfx

    Official

    Provides file paths to language-specific extension files for the code-testing pipeline.

    1k GitHub starsUsed in 2 repos~930 tokens
    Testing & QAAuto-check passed
  • Build Nitro Modules

    margelo/react-native-skills

    Builds and designs React Native Nitro Modules with Nitrogen, HybridObject TypeScript specs, Nitro View components, generated native implementations, zero-copy and native-state APIs, Swift/Kotlin/C++…

    175 GitHub stars~9.9k tokensUpdated 1 mo ago
    MobileAuto-check passed

More from dotnet/skills

All 91 skills in this repo
  • Official

    Resolves .NET runtime frames in Apple .ips crash logs to function names, source files and line numbers using dSYM symbols, atos and the Microsoft symbol server.

    5.6k GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Official

    Resolves native crash frames from .NET Android tombstones to function names, source files and line numbers using BuildIds, Microsoft's symbol server and llvm-symbolizer.

    5.6k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Official

    Scans C# and .NET code for about 50 performance anti-patterns and reports prioritized findings with concrete fixes, at a scan depth you choose.

    5.6k GitHub starsUsed in 3 repos~3.1k tokens
    Auto-check passed
  • Official

    Statically pairs source files with test files to list code that no test references, using Roslyn for C# or tree-sitter for many languages, with no build.

    5.6k GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check passed
  • Microbenchmarking

    dotnet/skills

    Official

    Activate this skill when BenchmarkDotNet (BDN) is involved in the task — creating, running, configuring, or reviewing BDN benchmarks.

    5.6k GitHub starsUsed in 3 repos~3.3k tokens
    Auto-check passed
  • Official

    Makes .NET projects compatible with Native AOT and trimming by resolving IL trim and AOT analyzer warnings through annotations rather than suppressions.

    5.6k GitHub starsUsed in 2 repos~4.2k tokens
    Auto-check passed

Categories

Questions about Grade Tests

What does Grade Tests do?

Assess a curated list of tests and produce a PR-ready table with a primary Pass, Failed, Uncertain, or Not applicable result plus A-F quality detail for every resolved test; Uncertain and Not…. Grade Tests is an agent skill from dotnet/skills, published by the product's own GitHub organization. Assess a curated list of tests and produce a PR-ready table with a primary Pass, Failed, Uncertain, or Not applicable result plus A-F quality detail for every resolved test; Uncertain and Not applicable omit the grade.

When should I use Grade Tests?

Grade Tests fits situations like: modified tests supplied as methods; A bounded PR diff; : suite-wide audits (use test-quality-auditor; test-anti-patterns).

How do I install Grade Tests in Claude Code?

Run `npx skills add dotnet/skills --skill grade-tests -a claude-code`. Or copy the skill folder (plugins/dotnet-test/skills/grade-tests in dotnet/skills) into .claude/skills/grade-tests in your project. Claude Code loads it when a task matches its description.

How do I install Grade Tests in Codex?

Run `npx skills add dotnet/skills --skill grade-tests -a codex`. Or copy the skill folder (plugins/dotnet-test/skills/grade-tests in dotnet/skills) into .agents/skills/grade-tests in your project. Codex loads it when a task matches its description.

Can I use Grade Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dotnet/skills --skill grade-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/grade-tests, .gemini/skills/grade-tests, .github/skills/grade-tests and .opencode/skills/grade-tests in your project.

What does Grade Tests need to run?

SKILL.md names no scripts, command-line tools or credentials: Grade Tests is instructions for the agent only.

Does Grade Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Grade Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Grade Tests use?

Grade Tests is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Grade Tests use?

About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Grade Tests?

Skills that share tags, products or a category with Grade Tests: Code Reviewer (alirezarezvani/claude-skills, 28k stars), Build Teaql App (teaql/teaql-agent-kit, 2.8k stars), Fory Version Bump (apache/fory, 4.6k stars) and Fory Performance Optimization (apache/fory, 4.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Grade Tests?

dotnet (a GitHub organization, an official publisher) maintains it in dotnet/skills, which has 5,568 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on October 7, 2026.

Source: dotnet/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.