Official agent skill

ONNX Runtime Test Runner

by microsoft in microsoft/onnxruntime

Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.

OfficialMITAuto-check passedTesting & QA

Install ONNX Runtime Test Runner

skills CLI
$ npx skills add microsoft/onnxruntime --skill ort-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/onnxruntime ort-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/ort-test .claude/skills/ort-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ort-test
GitHub stars
22k
Token cost
~1.8k tokens
SKILL.md length
804 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.

  • Works in 5 steps: Zero-match filter. A --gtest_filter that… → Stale binary from an incremental build.… → Checking the wrong artifact's freshness.… → …
  • Running a specific ONNX Runtime C++ test with a gtest filter
  • SKILL.md covers C++ tests, Python tests, Agent tips and False-green taxonomy — ways a…, plus 1 more section
  • Calls pytest

What it does

This skill explains how to run tests in the ONNX Runtime repository. C++ tests use Google Test through two executables: onnxruntime_test_all for the core framework, graph, optimizer and session, and onnxruntime_provider_test for operator and kernel tests across execution providers, selected with --gtest_filter. It warns about two same-named attention_op_test.cc files that test different operators, the ONNX-domain Attention operator and the contrib MultiHeadAttention and GroupQueryAttention ones.

Tests should always be run from the build output directory, which by default follows build/Platform/Config, may repeat the config name with Visual Studio generators and can be changed with --build_dir. A PowerShell search helps find a missing test binary, and build.sh or build.bat with --config Release --test runs everything after a successful build. Python tests use pytest by file, class, method or keyword, with unittest preferred.

When your agent uses it

  • Running a specific ONNX Runtime C++ test with a gtest filter
  • Debugging a failing ONNX Runtime test
  • Finding the test binary in a build output directory
  • Running ONNX Runtime Python tests with pytest

Example prompts

  • “Run the Conv3D provider tests in my Release build.”
  • “Which executable covers the optimizer tests, and how do I run a single case?”
  • “Run only the Python tests matching quantize under onnxruntime/test/python.”
  • “The attention test failed, so make sure I am running the right attention_op_test.cc.”

Requirements

  • A successful ONNX Runtime build that produced the test executables
  • pytest, for the Python tests

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Zero-match filter. A --gtest_filter that matches no tests still exits 0 (green).
  2. Stale binary from an incremental build. If the build did not actually recompile your
  3. Checking the wrong artifact's freshness. With a dlopen'd shared provider (e.g.
  4. A correct fallback path masks the intended path. A value-only assertion can pass via a
  5. Arch-portability false-green (verified on only one GPU arch). A CUDA kernel that

What it can do on your machine

Read from SKILL.md and the folder at commit 8420709. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ONNX Runtime Test Runner loads about 1.8k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 804 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/onnxruntime at commit 8420709, republished under its MIT licence (© microsoft). 804 words, ~1,827 tokens.

Download SKILL.mdSave it as .claude/skills/ort-test/SKILL.md (or your agent's skills folder).
name
ort-test
description
Run ONNX Runtime tests. Use this skill when asked to run tests, debug test failures, or find and execute specific test cases in ONNX Runtime.

Running ONNX Runtime Tests

ONNX Runtime uses Google Test for C++ and unittest (preferred) / pytest for Python.

C++ tests

Test executables
ExecutableWhat it tests
onnxruntime_test_allCore framework, graph, optimizer, session tests
onnxruntime_provider_testOperator/kernel tests (Conv, MatMul, etc.) across execution providers
Two attention_op_test.cc files — don't confuse them

There are two same-named files testing different operators. Both build into onnxruntime_provider_test:

PathOperatorgtest suite
test/providers/cpu/llm/attention_op_test.ccONNX-domain Attention (opset 23/24)AttentionTest.*
test/contrib_ops/attention_op_test.cccontrib MultiHeadAttention / GroupQueryAttentionContribOpAttentionTest.*

The MEA negative-offset regression tests (Attention_Causal_NonPadKVSeqLen_MEA_*, e.g. ..._MEA_NegOffset_ForceFlashDisabled_FP16_CUDA) live in the providers/cpu/llm file — the ONNX-domain op.

Use --gtest_filter to select specific tests:

bash
./onnxruntime_provider_test --gtest_filter="*Conv3D*"
Running tests

Always run from the build output directory — tests may fail to find dependencies otherwise.

bash
# Linux
cd build/Linux/Release
./onnxruntime_provider_test --gtest_filter="*TestName*"

# macOS
cd build/MacOS/Release
./onnxruntime_provider_test --gtest_filter="*TestName*"

# Windows
cd build\Windows\Release
.\onnxruntime_provider_test.exe --gtest_filter="*TestName*"

You can also run all tests via the build script (assumes a prior successful build):

bash
./build.sh --config Release --test
.\build.bat --config Release --test    # Windows
Locating the build output directory

The default path follows the pattern build/<Platform>/<Config>/ where Platform is Linux, MacOS, or Windows. With Visual Studio multi-config generators on Windows, the config may appear twice (e.g., build/Windows/Release/Release/). The path can also be customized via --build_dir.

If you can't find a test binary, search for it:

powershell
# Windows
Get-ChildItem -Path build -Recurse -Filter "onnxruntime_provider_test.exe" | Select-Object -ExpandProperty FullName

# Linux/macOS
find build -name "onnxruntime_provider_test" -type f

Python tests

Use pytest as the test runner:

bash
pytest onnxruntime/test/python/test_specific.py                          # entire file
pytest onnxruntime/test/python/test_specific.py::TestClass::test_method  # specific test
pytest -k "test_keyword" onnxruntime/test/python/                        # by keyword

Python test naming convention: test_<method>_<expected_behavior>_[when_<condition>]

Agent tips

  • Activate a Python virtual environment before running tests. See "Python > Virtual environment" in AGENTS.md.
  • Beware false-green results — a green run does not always prove anything. See the "False-green taxonomy" section below for the four ways a test can pass without testing your change.
  • Redirect test output to a file (e.g., > test_output.txt 2>&1) — output can be large.
  • For C++ tests, verify the build directory exists and a prior build completed before running.
  • Use --gtest_filter to run a targeted subset when the full suite takes too long.
  • Running WebGPU tests locally on Linux without a GPU — WebGPU op tests build into onnxruntime_provider_test and can run against a software Vulkan adapter (Mesa lavapipe). See the webgpu-local-testing skill.

False-green taxonomy — ways a test can "pass" without proving anything

A green result is not always a real pass. Watch for all five modes:

  1. Zero-match filter. A --gtest_filter that matches no tests still exits 0 (green). Confirm the [==========] N tests ran line is non-zero — a zero-match run prints 0 tests from 0 test suites. Many operator/kernel gtests run only in onnxruntime_provider_test (CI runs this), NOT onnxruntime_test_all; the wrong binary matches nothing and looks green.
  2. Stale binary from an incremental build. If the build did not actually recompile your change (e.g. a header not tracked by the compiler's depfile), the "passing" run executes the OLD code. A test that was failing cannot truly flip to passing without a real rebuild — treat an unexpected FAIL→PASS with suspicion and confirm the linked artifact's mtime advanced. CUDA/CUTLASS instance (nvcc depfiles don't track cutlass_fmha/*.h): see the cuda-cutlass-fmha-incremental-rebuild skill.
  3. Checking the wrong artifact's freshness. With a dlopen'd shared provider (e.g. libonnxruntime_providers_cuda.so), the test executable is NOT relinked when the provider recompiles — its mtime stays old while the .so advances. Verify the artifact that actually links your change, not the test exe. Detail: cuda-cutlass-fmha-incremental-rebuild skill.
  4. A correct fallback path masks the intended path. A value-only assertion can pass via a different, correct code path without ever exercising the one you meant to test (e.g. a test meant for MEA silently handled by the unfused fallback). Assert/verify which path ran, not just the output value — see "Verify which path/kernel actually executed" below.
  5. Arch-portability false-green (verified on only one GPU arch). A CUDA kernel that launches on a large-dynamic-smem arch (e.g. sm90/H100, ~227KB) can fail to launch on a smaller opt-in cap (sm86/89 ~99KB, sm80 ~163KB) with CUDA failure 1: invalid argument — and a path with no fallback (e.g. ORT's MEA) turns that into a hard error, not a silent degrade. So a green run on your local GPU can mask a launch failure on CI's arch. Verify arch-portability, or pick a config whose shared-memory footprint fits every target arch (e.g. a small head_size). Concrete instance: CUTLASS MEA head_size=512 FP16 exceeds sm86's smem opt-in cap and dies at launch — live bug #28388 (the cuda-attention-kernel-patterns skill §1 has the dispatch detail).
Show full SKILL.md (131 more words)Show less

Verify which path/kernel actually executed

Value equality alone does not prove the intended code path ran — a correct fallback can produce the right answer (false-green mode 4 above). When a test targets a specific kernel/path, confirm it actually dispatched there instead of trusting the output:

  • Enable verbose logging and check the dispatch log line. ORT attention logs one of these exact strings (core/providers/cuda/llm/attention.cc):
    • ONNX Attention: using Flash Attention (:1400)
    • ONNX Attention: using Memory Efficient Attention (:1451)
    • Attention: using unified unfused path (:1482) — note: no ONNX prefix and it reads "unified unfused path", not "Unfused".
  • Or force the path via the relevant env var / build config AND add a compile-time guard so the test SKIPs (not silently passes) when the target path is unavailable — e.g. SKIP_IF_MEA_NOT_COMPILED.

Operator-specific routing/forcing details: cuda-attention-kernel-patterns skill §1/§7.

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/ort-test of microsoft/onnxruntime.

Open the folder on GitHubat commit 8420709

Compare with similar skills

ONNX Runtime Test Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ONNX Runtime Test Runner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ONNX Runtime Test Runner this skillmicrosoft/onnxruntime22k—~1.8kAutomated safety check: PassMIT
Temporal Python Testingwshobson/agents40k11 repos~1.2kAutomated safety check: PassMIT
Squid Testing Pythoniusztinpaul/squid203—~1.3kAutomated safety check: PassApache-2.0
Flaky Test DetectorArabelaTso/Skills-4-SE253—~2kAutomated safety check: PassApache-2.0
Testing Pythonbenchflow-ai/skillsbench1.8k—~1.3kAutomated safety check: PassApache-2.0
Testing Patternssoftspark/ai-toolkit179—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Test Temporal workflows with pytest, time-skipping, and mocking strategies.

    40k GitHub starsUsed in 11 repos~1.2k tokens
    Testing & QAAuto-check passed
  • Squid Testing Python

    iusztinpaul/squid

    Write and evaluate effective Python tests using pytest. An agent skill from iusztinpaul/squid.

    203 GitHub stars~1.3k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Flaky Test Detector

    ArabelaTso/Skills-4-SE

    Identifies non-deterministic or unreliable tests through static code analysis and test result analysis.

    253 GitHub stars~2k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Testing Python

    benchflow-ai/skillsbench

    Write and evaluate effective Python tests using pytest. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.3k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Testing Patterns

    softspark/ai-toolkit

    Testing strategy: pyramid, AAA, mocks/fakes/stubs, flaky tests, coverage.

    179 GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Maintaining Python Tests

    PostHog/posthog-foss

    Official

    Maintains existing pytest and Django test suites without weakening correctness.

    721 GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check passed

More from microsoft/onnxruntime

All 14 skills in this repo
  • Official

    Finds and fixes out-of-range output writes in ONNX Runtime operator shape-inference functions where a getNumOutputs guard admits too few outputs.

    22k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Official

    Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

    22k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • ONNX Runtime Source Build

    microsoft/onnxruntime

    Official

    Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.

    22k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • ONNX Runtime CI Management

    microsoft/onnxruntime

    Official

    Triggers, re-runs and unblocks the CI checks on an ONNX Runtime pull request, after diagnosing whether a failure is transient or needs a code change.

    22k GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • ONNX Runtime Release Notes

    microsoft/onnxruntime

    Official

    Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component.

    22k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Official

    Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

    22k GitHub stars~2.9k tokensUpdated today
    Auto-check passed

Categories

Questions about ONNX Runtime Test Runner

What does ONNX Runtime Test Runner do?

Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance. This skill explains how to run tests in the ONNX Runtime repository. C++ tests use Google Test through two executables: onnxruntime_test_all for the core framework, graph, optimizer and session, and onnxruntime_provider_test for operator and kernel tests across execution providers, selected with --gtest_filter.

When should I use ONNX Runtime Test Runner?

ONNX Runtime Test Runner fits situations like: running a specific ONNX Runtime C++ test with a gtest filter; debugging a failing ONNX Runtime test; finding the test binary in a build output directory; running ONNX Runtime Python tests with pytest.

How do I install ONNX Runtime Test Runner in Claude Code?

Run `npx skills add microsoft/onnxruntime --skill ort-test -a claude-code`. Or copy the skill folder (.github/skills/ort-test in microsoft/onnxruntime) into .claude/skills/ort-test in your project. Claude Code loads it when a task matches its description.

How do I install ONNX Runtime Test Runner in Codex?

Run `npx skills add microsoft/onnxruntime --skill ort-test -a codex`. Or copy the skill folder (.github/skills/ort-test in microsoft/onnxruntime) into .agents/skills/ort-test in your project. Codex loads it when a task matches its description.

Can I use ONNX Runtime Test Runner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/onnxruntime --skill ort-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ort-test, .gemini/skills/ort-test, .github/skills/ort-test and .opencode/skills/ort-test in your project.

What does ONNX Runtime Test Runner need to run?

Going by SKILL.md and its folder, ONNX Runtime Test Runner needs the command-line tools its instructions call (pytest). Our summary lists: A successful ONNX Runtime build that produced the test executables; pytest, for the Python tests.

Does ONNX Runtime Test Runner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is ONNX Runtime Test Runner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does ONNX Runtime Test Runner use?

ONNX Runtime Test Runner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ONNX Runtime Test Runner use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ONNX Runtime Test Runner?

Skills that share tags, products or a category with ONNX Runtime Test Runner: Temporal Python Testing (wshobson/agents, 40k stars), Squid Testing Python (iusztinpaul/squid, 203 stars), Flaky Test Detector (ArabelaTso/Skills-4-SE, 253 stars) and Testing Python (benchflow-ai/skillsbench, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ONNX Runtime Test Runner?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/onnxruntime, which has 22,029 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 7, 2026.

Source: microsoft/onnxruntime on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.