Build Teaql App
teaql/teaql-agent-kit
Build or change a TeaQL application in Java, Rust, Go, Swift, Python, C/.NET, or TypeScript, including Kotlin/JVM applications that consume Java-generated libraries.
Grade a curated list of individual tests for readiness, A-F quality, and concrete improvements.
$ npx skills add microsoft/testfx --skill grade-tests -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/testfx grade-tests --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/testfx.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/grade-tests .claude/skills/grade-tests && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "grade-tests" agent skill from https://github.com/microsoft/testfx/tree/main/.agents/skills/grade-tests into .claude/skills/grade-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grade-tests", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/testfx/tree/main/.agents/skills/grade-testsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/testfx --skill grade-tests -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/testfx grade-tests --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/testfx.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/grade-tests .agents/skills/grade-tests && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "grade-tests" agent skill from https://github.com/microsoft/testfx/tree/main/.agents/skills/grade-tests into .agents/skills/grade-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grade-tests", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/testfx --skill grade-tests -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/testfx grade-tests --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/testfx.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/grade-tests .cursor/skills/grade-tests && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "grade-tests" agent skill from https://github.com/microsoft/testfx/tree/main/.agents/skills/grade-tests into .cursor/skills/grade-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grade-tests", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/testfx.git --path .agents/skills/grade-tests--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/testfx --skill grade-tests -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/testfx grade-tests --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/testfx.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/grade-tests .gemini/skills/grade-tests && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "grade-tests" agent skill from https://github.com/microsoft/testfx/tree/main/.agents/skills/grade-tests into .gemini/skills/grade-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grade-tests", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/testfx grade-testsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/testfx --skill grade-tests -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/testfx.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/grade-tests .github/skills/grade-tests && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "grade-tests" agent skill from https://github.com/microsoft/testfx/tree/main/.agents/skills/grade-tests into .github/skills/grade-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grade-tests", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/testfx --skill grade-tests -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/testfx grade-tests --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/testfx.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/grade-tests .opencode/skills/grade-tests && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "grade-tests" agent skill from https://github.com/microsoft/testfx/tree/main/.agents/skills/grade-tests into .opencode/skills/grade-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grade-tests", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
grade-testsGrade a curated list of individual tests for readiness, A-F quality, and concrete improvements.
Grade Tests is an agent skill from microsoft/testfx, published by the product's own GitHub organization. Grade a curated list of individual tests for readiness, A-F quality, and concrete improvements. ALWAYS USE FOR: grade tests, review only a named test, per-test readiness decisions, or quality bands for supplied methods, bodies, file spans, or bounded PR diffs, including existing tests. Produce a PR-ready Pass, Failed, Uncertain, or Not applicable table; unresolved or empty scopes omit the grade. Compose read-only per-test mutation evidence when available. Polyglot: .NET, Python, TS/JS, Java, Go, Ruby, Rust…
Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Mobile, covering Android development and iOS development. It works with Python, .NET, TypeScript and C++. The repository describes itself as: This repository holds the source code of Microsoft.Testing.Platform (MTP), a lightweight alternative to VSTest, as well as MSTest adapter and framework. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 7226b0c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Grade Tests loads about 6.8k tokens when it runs. Until then it costs about 169 tokens; SKILL.md has 3,471 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from microsoft/testfx at commit 7226b0c, republished under its MIT licence (© microsoft). 3,471 words, ~6,778 tokens.
.claude/skills/grade-tests/SKILL.md (or your agent's skills folder).Assess a curated list of test methods and produce a compact, PR-comment-friendly report. The primary result is one of Pass, Failed, Uncertain, or Not applicable; an A-F quality grade remains secondary diagnostic information. The skill does not discover tests on its own — the caller (typically a PR automation workflow or a human reviewer holding a specific list) provides the tests or a bounded diff to assess.
After Step 0 admits a bounded scope, enforce these grading invariants:
test-gap-analysis by name once before
scoring, in per-test-read-only caller context. Read its owned composition
reference; do not compute mutation evidence from this grading rubric or run
its standalone workflow. Report N/A / unverified only when the dependency,
reference, or required context is actually unavailable.Language-specific guidance: If the caller supplies the matching bundled extension file path, read it directly. Otherwise call
test-analysis-extensionsto discover available extension files, then read the file matching the target codebase's language and framework (e.g.,extensions/dotnet.md,extensions/python.md,extensions/typescript.md,extensions/go.md). You MUST read the relevant extension file before scoring assertions or anti-patterns, because assertion APIs and idiomatic patterns differ significantly across frameworks.
PR reviewers need a simple answer to does this test need follow-up? The four-state result provides that decision; the existing A-F rubric explains its quality and severity.
test-anti-patterns (pragmatic) or test-smell-detection (formal) and
let the test-engineer agent orchestrate its internal quality specialist.test-engineer
(any language) or writing-mstest-tests (MSTest specifically).coverage-analysis or crap-score (.NET only).| Input | Required | Description |
|---|---|---|
| Test methods | Yes | A scope to grade. Provide one of: (a) an explicit list of test method names (fully-qualified, e.g. Namespace.ClassName.TestMethodName); (b) one or more file paths plus an explicit instruction to grade every test declared in those files; or (c) a diff hunk / PR identifier whose changed tests should be graded. File paths are recommended but optional when method names are unambiguous in the workspace. Ambiguous requests like "grade my tests" with no scope are rejected up-front (see Step 0); this skill is for curated input and does not auto-grade an entire workspace. |
| Test bodies / spans | Recommended | The exact source lines for each test method. If omitted, read them from the listed files. |
| Production code | No | The code under test, for judging whether assertions cover the claimed behavior. When unavailable, mark the mutation assessment N/A / unverified rather than guessing or deducting. |
| Language reference | No | A host-supplied path to the matching bundled test-analysis-extensions file. Read it directly instead of invoking its reference-only loader; do not substitute unverified framework guidance. |
| Diff context | No | When grading PR changes, the unified diff for each test method helps focus on what actually changed. |
Before doing anything else, check that the caller provided one of:
OrderTests.cs"), orIf the request is ambiguous (e.g., "Grade my tests", "Are these tests
any good?" with no scope, "Review the test suite"), do not load
extensions, do not read files, and do not grade anything. Reply with a
short message asking the caller to provide an explicit list / file(s) /
diff, and optionally point them at the test-engineer agent or
test-anti-patterns skill for full-suite analysis. Stop there.
If a valid bounded scope resolves to zero eligible tests, return Not applicable with a short explanation and no invented rows.
Identify the target codebase's language and test framework from the file
extensions and the test method markers in the provided list. Call the
test-analysis-extensions skill unless the caller already supplied the matching
bundled extension file path. In either case, read that extension file (e.g.,
extensions/dotnet.md for MSTest/xUnit/NUnit/TUnit, extensions/python.md
for pytest, extensions/typescript.md for Jest/Vitest, extensions/go.md
for the standard testing package). If the input contains tests from
multiple languages, load each relevant extension and grade each test using
its language's conventions.
For each entry in the input list:
Uncertain — method not found with no quality grade and continue. Never
invent a body to grade. A missing requested method requires human review;
it is not the same as a valid scope containing no tests.Composition checkpoint: for resolved tests with available production
context, load test-gap-analysis now, once for the batch, with
per-test-read-only assessment context. Complete its owned reference assessment
before Step 3. Do not skip this load just because a body-level weakness already
seems obvious; a locally invented mutation explanation is not composition.
Keep grading read-only: no build/test runs, mutation execution, file edits, tool installation, broad suite discovery, or agent delegation. Resolve only the supplied tests, their relevant fixtures/helpers, and the production call chain needed for their claims.
Use the inline test-gap-analysis assessment from Step 2's checkpoint;
do not load it a second time. Supply each test's
identifier/body, relevant setup/helpers, claimed behavior, assertion semantics,
and available source. Its composition dispatch loads the owned read-only
reference rather than its standalone baseline/verification workflow. Consume
its per-test evidence; do not duplicate its mutation catalog here or invoke an
audit/generation agent.
Convey mode and inputs as assessment context using the host's supported caller
instructions. If the loader accepts only a skill name, load test-gap-analysis
by name only; do not invent tool arguments or a mode-specific skill name.
If the skill/reference or production context is unavailable, record
Pseudo-mutation: N/A / unverified — <reason> and continue normal body-level
grading. This is not a grade deduction or, by itself, an Uncertain result.
Do not search installation directories or substitute a mutation runner.
Assess only what each test claims: do not borrow another test's assertions, or demand unrelated branches, outputs, or scenarios. An observable survivor can support an existing Assertion strength category when it proves that the test does not verify its claimed outcome; do not introduce mutation points, weights, ceilings, or an automatic survivor penalty. Apply the existing rubric normally, including weaknesses it classifies in both Assertion and Anti-pattern dimensions; do not add another deduction for the same mutation evidence.
Start every test at grade A (score band 90–100), then apply deductions strictly for observable issues in the captured body. Do not deduct for hypothetical concerns (e.g., "could have more negative assertions") unless the production code clearly demands them and the production code is available.
When production code is unavailable, grade observable issues in the test body
normally, but do not infer missing behaviors or deduct for them. State
Production-dependent behavior coverage: Unverified once in the summary so the
reader can distinguish test-body findings from claims that require source code.
Compute three sub-grades (each A–F) that together drive the overall grade.
Read the loaded language extension's assertion API list and classify every assertion in the test body. Score from highest to lowest:
| Sub-grade | Pattern |
|---|---|
| A | At least one meaningful value assertion (equality / structural / exception / state) plus, where appropriate, additional checks (negative, type, collection contents). Mock-call verifications (Verify, toHaveBeenCalledWith, Should -Invoke) and bare assertion forms (pytest assert, Go if got != want { t.Errorf(...) }, Rust assert!()) count as real assertions. |
| B | One clear meaningful assertion that verifies the behavior under test. |
| C | Only trivial assertions (single IsNotNull / toBeDefined / assert x is not None), or assertions that leave a meaningful part of the test's claimed result unchecked. A focused single-field claim does not require unrelated fields. |
| D | One self-referential / tautological assertion (Assert.AreEqual(x, x), assert dto.name == dto.name, round-trip identity without a non-trivial input), or broad exception assertions (Assert.ThrowsException<Exception>). |
| F | No assertions at all; all assertions are always-true literals (Assert.IsTrue(true), assert True, expect(true).toBe(true)) — these verify nothing and are equivalent to having no assertions; or all assertions are silently un-awaited (e.g., expect(promise).resolves.toBe(x) without await/return, async TUnit/xUnit Assert.ThrowsAsync without await, pytest-asyncio with un-awaited coroutine). |
Exception and error-path tests (Assert.ThrowsException<T>, constrained
pytest.raises, expect(fn).toThrow, assertThrows, #[should_panic],
Should -Throw, EXPECT_THROW, or Go code that verifies an expected non-nil
error) are complete on their own. Give Assertion strength A when the test
checks the exact promised error condition for its stated scope. Do not deduct
for having only that assertion, and do not require an error-message assertion
unless the message is part of the documented contract. A Go happy-path test
that only checks err == nil while discarding a meaningful returned value is
still C because it does not verify the successful result.
| Sub-grade | Pattern |
|---|---|
| A | Clear Arrange-Act-Assert (or Given-When-Then) separation. Single behavior under test. Body under ~30 lines. Setup uses framework conventions. |
| B | One mild structural issue (slightly long body, missing blank lines between phases) but intent is clear. |
| C | Multiple behaviors mixed in one test, or AAA phases interleaved enough to slow comprehension. |
| D | Conditional logic in the test (if/switch driving assertions) — except for idiomatic Go/Rust table-driven sub-test loops; or test relies on previous test state (ordering dependency). |
| F | Test exceeds ~60 lines and verifies multiple unrelated behaviors; or shares mutable state with other tests through statics/globals without reset. |
Scan against the catalog below. The Anti-pattern sub-grade is computed in two passes and combined deterministically:
The final Anti-pattern sub-grade is the worse of the two passes
(i.e., min(hard_ceiling, A − medium_count)). Low findings never
affect the grade — mention them in the note only.
Examples (Critical/High and Medium counts → Anti-pattern sub-grade):
min(C, A − 2 = C) = C; a third Medium tips to D)Critical (drop straight to F or D)
try { … } catch { } (.NET), bare except: pass
(Python), try { … } catch (e) {} (JS/TS/Java), defer recover()
without re-panic (Go), rescue StandardError with no assertion (Ruby),
empty catch (Kotlin/Swift) → FAssert.Fail(ex.Message) instead of
Assert.ThrowsException) → DAssert.IsTrue(true), assert True,
expect(true).toBe(true)) → F (verifies nothing; also drives
Assertion sub-grade to F)Assert.AreEqual(x, x), assert dto.name == dto.name) → DHigh (drop one or two sub-grades)
Thread.Sleep, Task.Delay,
time.sleep, setTimeout-based wait, Thread.sleep, time.Sleep,
sleep, std::thread::sleep, Start-Sleep,
std::this_thread::sleep_for (in a unit test) → DDateTime.Now, datetime.now(), Date.now(),
System.currentTimeMillis(), time.Now(), Time.now,
Instant::now(), Get-Date, system_clock::now) → DC:\…, /tmp/…, network hosts) → DAssert.ThrowsException<Exception>,
pytest.raises(Exception), expect(fn).toThrow(Error) without matcher,
#[should_panic] without expected = "…", Should -Throw without
-ExpectedMessage, EXPECT_ANY_THROW) → CMedium (drop one sub-grade)
Test1, TestMethod, test, single-word name that says
nothing about scenario or expected outcome (judge against the language
extension's convention) → drop one sub-grade42, "foo", 0x1234 in arrange/assert
without naming or comment → drop one sub-gradeLow (note only, no deduction)
Console.WriteLine,
print, console.log, System.out.println, fmt.Println, puts,
dbg!, Write-Host, std::cout); inconsistent naming versus siblings;
leftover TODO comments. Mention in the note column but do not deduct.Convert sub-grades to numeric points: A=4, B=3, C=2, D=1, F=0.
0.45 × Assertion + 0.30 × Anti-pattern + 0.25 × StructureReport the letter grade and the score band (not a single 0–100 number). False precision invites bikeshedding; bands keep the conversation focused on the rubric.
The grade summarizes strength; the result says whether follow-up exists. An actionable improvement is an evidence-backed change to the test, setup, or fixtures. Assign exactly one:
Do not derive status from grade: a complete focused test can be B / Pass, while debug output can make an otherwise excellent test A / Failed. Use Uncertain for an unresolved body, unsupported construct, or essential missing contract—not merely absent production code. A definite finding wins over uncertainty.
Use one sentence (target ≤ 120 characters) for the most important reason:
No issues found., Only checks IsNotNull; receipt contents are unverified., or
Method body could not be resolved; human review is required. Do not invent a
weakness to justify a grade or Failed result.
Keep the action in a separate How to improve field. For each Failed test,
name the smallest useful input, assertion, or fixture change and its expected
outcome, grounded in the body, source, or an explicit contract. For example,
Replace self-comparison with Assert.AreEqual(60m, account.Balance)., not
Improve assertions; Remove Console.WriteLine after Deposit(25m)., not
Clean up. Prioritize the highest-impact distinct finding, and include other
actionable findings only when they require a different change.
For a behavioral gap, use the distinguishing witness and original/mutant
observations from the shared assessment; check the expected result against
the unmodified source. If essential context is missing, name the evidence
needed instead of inventing an expected value. Pass gets None; Uncertain
gets a concrete evidence-resolution step, not a speculative test rewrite.
A rubric-only deduction is not proof of a behavioral gap or an actionable
improvement: a focused B / Pass may need no change. A / Failed still
needs its concrete action, such as removing debug output.
Produce two sections.
Begin with **Result: <Pass|Failed|Uncertain|Not applicable>**, then give result
counts and the highest-priority action. Aggregate using
Failed → Uncertain → Pass → Not applicable. For Not applicable, explain the
empty scope and omit the table.
| Test | Result | Quality | Notes | How to improve |
|------|--------|---------|-------|----------------|
| `Namespace.ClassName.Test_Method_Condition_Expected` | Pass | B (80–89) | One complete value assertion. | None |
| `Namespace.ClassName.Withdraw_SufficientFunds` | Failed | D (60–69) | Balance is compared with itself. | Replace self-comparison with `Assert.AreEqual(60m, account.Balance)` after withdrawing 40m from 100m. |
| `Namespace.ClassName.Test_Missing` | Uncertain | — | Method body could not be resolved; human review is required. | Supply the method body and its referenced fixture. |Keep these two report sections and the original Test/Result/Quality/Notes fields. When mutation evidence explains a finding or the caller requests detail, append a compact per-test Pseudo-mutation evidence block inside the per-test section: change, witness, original/mutant observations, relevant assertion, and classification. Static results are Likely killed (inferred) or Candidate survivor (unverified), never executed Killed/Survived or empirical killed/total counts. State missing-context N/A / unverified once per shared limitation. Do not repeat the improvement table in prose.
Caps and ordering:
<details> block.(new) or
(modified) marker.If multiple languages are present, produce one table per language and prefix each section with the language name and framework.
Uncertain — method not found).Assert.IsTrue(result.IsValid))
are not classified as always-true; only literal true/false constants are.assert, Go if got != want { t.Errorf(...) },
JS/TS expect(mock).toHaveBeenCalledWith(...).resolves/rejects/ThrowsAsync,
pytest-asyncio without await) drop the Assertion sub-grade to F.| Pitfall | Solution |
|---|---|
| Grading every test in the workspace when no list is provided | Ask the caller for the explicit list; this skill is for curated input. |
| Inflating deductions to justify the grade | Start at A; deduct only for observable issues. |
| Penalizing exception tests for low assertion count | Exception assertions are complete on their own. |
Downgrading a focused Go error-path test because it checks only err != nil | Expected-error existence is the observable contract for that scope; keep it at A unless the production contract requires a specific error identity or message. |
Treating IsNotNull before a value assertion as trivial | Only flag when the null check is the only assertion. |
| Treating any Boolean assertion as effectively assertion-free | Only always-true literals (Assert.IsTrue(true), assert True) are; meaningful Assert.IsTrue(result.IsValid) is a real assertion. |
| Flagging Go/Rust table-driven loops as conditional logic | They are idiomatic; do not deduct. |
Treating pytest bare assert or Go if got != want { t.Error… } as missing-framework | Both are canonical; count in the correct assertion category. |
| Penalizing tests when production code is unavailable | Mark concerns about uncovered behaviors as Unverified and do not deduct. |
| Using a fake-precise score (e.g., 87/100) | Use the score band only — 90–100, 80–89, 70–79, 60–69, 0–59. |
| Spilling a 500-row table into a PR comment | Apply the row cap from Step 6; collapse extras into <details>. |
| Re-reporting an existing finding three times under different categories | Pick the most fitting category and report once. |
| Giving a weak test credit for a sibling's assertions | Use only the current test and helpers/fixtures it executes. |
| Turning pseudo-mutation composition into a suite audit | Pass explicit per-test-read-only mode; no runs, edits, broad discovery, or agent recursion. |
| Inventing weaknesses for A-grade tests to make the note "balanced" | If a test is clean, the note may simply read No issues found. |
| Mapping status from grade or comments | Fail only for actionable improvements; a B can Pass and an A can Fail. |
| Confusing Uncertain and Not applicable | Evidence gaps are Uncertain; a valid empty scope is Not applicable. |
© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/grade-tests of microsoft/testfx.
Open the folder on GitHubat commit 7226b0c
We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in microsoft/testfx, which our catalogue first saw on October 10, 2026.
Grade Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Grade Tests this skillmicrosoft/testfx | 1k | 2 repos | ~6.8k | Automated safety check: Pass | MIT | |
| Build Teaql Appteaql/teaql-agent-kit | 2.8k | — | ~4.6k | Automated safety check: Pass | MIT | |
| Code Revieweralirezarezvani/claude-skills | 28k | 1 repos | ~1.6k | Automated safety check: Pass | MIT | |
| Fory Version Bumpapache/fory | 4.6k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Fory Performance Optimizationapache/fory | 4.6k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Build Nitro Modulesmargelo/react-native-skills | 175 | — | ~9.9k | Automated safety check: Pass | MIT |
teaql/teaql-agent-kit
Build or change a TeaQL application in Java, Rust, Go, Swift, Python, C/.NET, or TypeScript, including Kotlin/JVM applications that consume Java-generated libraries.
alirezarezvani/claude-skills
Code review automation for TypeScript, JavaScript, Python, Go, Swift, Kotlin, C, .NET, Java, C, C++, Rust, Ruby, PHP, and Dart/Flutter.
apache/fory
Bump Apache Fory release or post-release development versions across Java, Kotlin, Scala, Python, Rust, Go, C++, C, Dart, JavaScript, Swift, integration tests, examples, and source docs.
apache/fory
Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala).
margelo/react-native-skills
Builds and designs React Native Nitro Modules with Nitrogen, HybridObject TypeScript specs, Nitro View components, generated native implementations, zero-copy and native-state APIs, Swift/Kotlin/C++…
corvus-dotnet/Corvus.JsonSchema
Work on the Java port of the V5 standalone schema evaluator (src-java/corvus-json-schema, Maven artifact io.github.corvus-dotnet:corvus-json-schema): loader, compiler, the ASM bytecode generator…
microsoft/testfx
MANDATORY for static source-to-test pairing: find or list source files/modules without corresponding tests, or suggest test locations from repository structure.
microsoft/testfx
Activation requires either supplied .NET coverage reports/percentages/line, branch, or condition metrics, or an explicit request to collect .NET coverage for analysis.
microsoft/testfx
Safely refactors C/.NET code without changing behavior. An agent skill from microsoft/testfx.
microsoft/testfx
Identify a .NET project's test platform, framework, command mode, and SDK-style vs classic project system.
microsoft/testfx
Validate TestFx shipping paths and capable CI execution using packed consumers, package/cache provenance, exact exits and artifacts, and selected-versus-executed test evidence.
microsoft/testfx
ALWAYS USE when asked to fix, rewrite, update, improve, modernize, show corrected code for, or explain existing MSTest tests or MSTest-specific configuration.
Categories
Grade a curated list of individual tests for readiness, A-F quality, and concrete improvements. Grade Tests is an agent skill from microsoft/testfx, published by the product's own GitHub organization. Grade a curated list of individual tests for readiness, A-F quality, and concrete improvements.
Grade Tests fits situations like: review only a named test; per-test readiness decisions; quality bands for supplied methods; bounded PR diffs.
Run `npx skills add microsoft/testfx --skill grade-tests -a claude-code`. Or copy the skill folder (.agents/skills/grade-tests in microsoft/testfx) into .claude/skills/grade-tests in your project. Claude Code loads it when a task matches its description.
Run `npx skills add microsoft/testfx --skill grade-tests -a codex`. Or copy the skill folder (.agents/skills/grade-tests in microsoft/testfx) into .agents/skills/grade-tests in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/testfx --skill grade-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/grade-tests, .gemini/skills/grade-tests, .github/skills/grade-tests and .opencode/skills/grade-tests in your project.
SKILL.md names no scripts, command-line tools or credentials: Grade Tests is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Grade Tests is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Grade Tests: Build Teaql App (teaql/teaql-agent-kit, 2.8k stars), Code Reviewer (alirezarezvani/claude-skills, 28k stars), Fory Version Bump (apache/fory, 4.6k stars) and Fory Performance Optimization (apache/fory, 4.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/testfx, which has 1,047 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 9, 2026.
Source: microsoft/testfx on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.