Debug Failing GPU
facebookexperimental/triton
Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.
Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.
$ npx skills add microsoft/onnxruntime --skill ort-transformers-gpu-pytest -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/onnxruntime ort-transformers-gpu-pytest --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/ort-transformers-gpu-pytest .claude/skills/ort-transformers-gpu-pytest && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ort-transformers-gpu-pytest" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-transformers-gpu-pytest into .claude/skills/ort-transformers-gpu-pytest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-transformers-gpu-pytest", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-transformers-gpu-pytestType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/onnxruntime --skill ort-transformers-gpu-pytest -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/onnxruntime ort-transformers-gpu-pytest --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.github/skills/ort-transformers-gpu-pytest .agents/skills/ort-transformers-gpu-pytest && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ort-transformers-gpu-pytest" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-transformers-gpu-pytest into .agents/skills/ort-transformers-gpu-pytest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-transformers-gpu-pytest", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/onnxruntime --skill ort-transformers-gpu-pytest -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/onnxruntime ort-transformers-gpu-pytest --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.github/skills/ort-transformers-gpu-pytest .cursor/skills/ort-transformers-gpu-pytest && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ort-transformers-gpu-pytest" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-transformers-gpu-pytest into .cursor/skills/ort-transformers-gpu-pytest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-transformers-gpu-pytest", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/onnxruntime.git --path .github/skills/ort-transformers-gpu-pytest--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/onnxruntime --skill ort-transformers-gpu-pytest -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/onnxruntime ort-transformers-gpu-pytest --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.github/skills/ort-transformers-gpu-pytest .gemini/skills/ort-transformers-gpu-pytest && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ort-transformers-gpu-pytest" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-transformers-gpu-pytest into .gemini/skills/ort-transformers-gpu-pytest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-transformers-gpu-pytest", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/onnxruntime ort-transformers-gpu-pytestInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/onnxruntime --skill ort-transformers-gpu-pytest -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .github/skills && cp -r skills-src/.github/skills/ort-transformers-gpu-pytest .github/skills/ort-transformers-gpu-pytest && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ort-transformers-gpu-pytest" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-transformers-gpu-pytest into .github/skills/ort-transformers-gpu-pytest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-transformers-gpu-pytest", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/onnxruntime --skill ort-transformers-gpu-pytest -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/onnxruntime ort-transformers-gpu-pytest --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.github/skills/ort-transformers-gpu-pytest .opencode/skills/ort-transformers-gpu-pytest && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ort-transformers-gpu-pytest" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-transformers-gpu-pytest into .opencode/skills/ort-transformers-gpu-pytest/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-transformers-gpu-pytest", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ort-transformers-gpu-pytestRuns the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.
The skill specializes the general `ort-test` skill for the Python transformers tests under `onnxruntime/test/python/transformers/` on a machine that also has `torch` installed, and it lists gotchas that would each cost an hour to rediscover. The first is never to run pytest from the repository root: the source folder `onnxruntime/` shadows the installed wheel and produces `ModuleNotFoundError: No module named 'onnxruntime.capi'`. The fix is a fresh private directory from `mktemp -d`, the test helper directory on `PYTHONPATH` and the test file given by absolute path.
A bare `/tmp` is avoided because pytest puts the working directory on the path and Python would import a `sitecustomize.py` planted by another user on a shared box. The second gotcha is that torch's bundled CUDA and cuDNN libraries can shadow the ones ORT was built with, which sends the SDPA decode tier to `MATH` instead of `CUDNN_FLASH_ATTENTION`, so the skill pins the libraries with `LD_PRELOAD`. The aim throughout is dispatch-verified runs, where a test is shown to have exercised the cuDNN decode tier and not skipped or fallen back silently.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8420709. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonpytestFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
ONNX Runtime GPU Transformers Tests loads about 2.9k tokens when it runs. Until then it costs about 120 tokens; SKILL.md has 1,163 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from microsoft/onnxruntime at commit 8420709, republished under its MIT licence (© microsoft). 1,163 words, ~2,900 tokens.
.claude/skills/ort-transformers-gpu-pytest/SKILL.md (or your agent's skills folder).Reusable, hard-won knowledge for running the Python transformers tests under
onnxruntime/test/python/transformers/ against a GPU-built wheel, on a box where
torch is also installed. Three gotchas below will each cost you an hour if
rediscovered. See the ort-test skill for the general test taxonomy and the
false-green modes; this skill is the GPU + transformers + SDPA-dispatch specialization.
Symptom:
ModuleNotFoundError: No module named 'onnxruntime.capi'
# or an AttributeError deep inside onnxruntime importCause: the repo root contains a source directory ./onnxruntime/ (the C++/Python
source tree). When pytest runs with the repo root on sys.path[0], import onnxruntime
resolves to that source package — which has no compiled capi extension — instead
of the installed wheel in your venv's site-packages. The source dir shadows the
wheel.
Fix: run pytest from a neutral, private working directory — a fresh
mktemp -d, not the repo root and not a bare shared /tmp — so the shadowing source
dir is not on the path, and point at the test file by absolute path. Put the
transformers test-helper dir on PYTHONPATH so shared helpers still import.
WORKDIR=$(mktemp -d); cd "$WORKDIR" # NEUTRAL + PRIVATE cwd — NOT repo root, NOT /tmp
export PYTHONPATH=/abs/repo/onnxruntime/test/python/transformers
python -m pytest /abs/repo/onnxruntime/test/python/transformers/<file>.py -vWhy mktemp -d and not a bare cd /tmp: pytest prepends the cwd to sys.path, and
Python auto-imports sitecustomize.py/usercustomize.py from it at startup — so on a
shared box a co-tenant's planted /tmp/sitecustomize.py would execute as arbitrary code
in your test process. A fresh per-run mktemp -d (private, 0700) keeps the neutral-cwd
source-shadow protection while removing that injection vector.
Do not cd into the repo and run pytest onnxruntime/test/... — that reintroduces
the shadowing. (This is the Python analogue of the C++ "run from the build output dir"
rule in ort-test.)
Symptom: any of —
MATH instead of CUDNN_FLASH_ATTENTION
(wrong-version cuDNN loaded), orlibcudnn.so.9: cannot open shared object file / undefined symbol / cuDNN version
mismatch errors at first CUDA op, orCause: a pip-installed torch ships its own bundled CUDA runtime + cuDNN
(e.g. cu124 → CUDA 12.4 / cuDNN 9.1) under site-packages/nvidia/*/lib. If ORT was
built against a different CUDA/cuDNN (e.g. CUDA 12.9 / cuDNN 9.8), whichever set the
dynamic loader resolves first wins. With torch imported (or its libs on the path),
torch's older libs can shadow the ones ORT dlopens → wrong-version dispatch or load
failure.
Fix: LD_PRELOAD the system CUDA runtime + cuDNN that ORT was built against so
they are loaded first, and add their dirs to LD_LIBRARY_PATH. Activate the venv that
has the ORT wheel. Concrete form used successfully (CUDA 12.9 + cuDNN 9.8; substitute
your absolute lib paths):
WORKDIR=$(mktemp -d); cd "$WORKDIR" # neutral + private (see §1)
source /abs/repo/.venv/bin/activate
export LD_PRELOAD=/abs/cuda12.9/lib64/libcudart.so.12:/abs/cudnn9.8/lib/libcudnn.so.9
export LD_LIBRARY_PATH=/abs/cuda12.9/lib64:/abs/cudnn9.8/lib
export PYTHONPATH=/abs/repo/onnxruntime/test/python/transformers
python -m pytest /abs/repo/onnxruntime/test/python/transformers/<file>.py -vNotes:
libcudart.so.12 is correct for both CUDA 12.4 and 12.9 (SONAME is major-only) —
pinning the 12.9 file forces the right minor.torch.cuda still works fine under this preload — the bf16 IO-binding path that uses
torch tensors + .data_ptr() runs correctly.LD_LIBRARY_PATH can still let
a transitive dependency resolve against torch's copy.A numerically-correct result does not prove the cuDNN SDPA path ran — the kernel has
a MATH fallback that produces the same answer (false-green mode 4 in ort-test). To
prove the tier dispatched, observe ORT's routing rather than probing a version.
ORT's ONNX-domain Attention kernel emits a debug line when
ORT_ENABLE_ATTENTION_KERNEL_DEBUG_INFO=1 is set before the InferenceSession is
created (the option is read once at session creation). AttentionKernelDebugInfo::Print
emits a token of the form:
SdpaKernel=CUDNN_FLASH_ATTENTION # or =MATH, =FLASH_ATTENTION, =EFFICIENT_ATTENTION (non-exhaustive)Capture stdout across a single run() and parse it. Capture at the file-descriptor
level, not contextlib.redirect_stdout: the SdpaKernel= line is written to native
fd-1 from C++, which Python-level stdout redirection never intercepts — you would get
dispatched=None, a silent false-negative. Mirror ORT's own _CaptureStdout
(os.dup2 fd-1 to a temp file, run, restore, read it back):
os.environ["ORT_ENABLE_ATTENTION_KERNEL_DEBUG_INFO"] = "1" # BEFORE InferenceSession()
# FD-level capture (see onnxruntime's _CaptureStdout for the exact idiom):
saved_fd = os.dup(1)
tmp = tempfile.TemporaryFile()
os.dup2(tmp.fileno(), 1) # redirect native fd-1
try:
# ... create session, run once ...
finally:
os.dup2(saved_fd, 1) # restore fd-1
os.close(saved_fd)
tmp.seek(0)
captured_text = tmp.read().decode()
m = re.search(r"SdpaKernel=(?P<kernel>[A-Z_]+)", captured_text)
dispatched = m.group("kernel") if m else None
assert dispatched == "CUDNN_FLASH_ATTENTION"Caveat: re.search returns only the first SdpaKernel= token — correct for the
single-node decode probe here. For a graph with multiple attention nodes use
re.findall and check every token, or a later node's MATH fallback is masked by an
earlier cuDNN hit.
Prefer this over reading torch.backends.cudnn.version() or any library-version check:
a version probe reads torch's cuDNN, not the cuDNN ORT actually loaded/dispatched —
that mismatch is a real trustworthiness bug. Observe-dispatch reads ORT's own routing
decision, so it is correct across cuDNN versions with no hard-coded version table.
ORT_TEST_REQUIRE_CUDNN_SDPAGating decode tests by observed dispatch has a failure mode: if the tier silently regresses (stops selecting cuDNN), the observation returns "not dispatched" and every decode test skips green, hiding the regression as all-green skips.
Close the hole with an env-gated canary: ORT_TEST_REQUIRE_CUDNN_SDPA=1. When set, the
dispatch assertion becomes non-skippable — a MATH fallback / non-dispatch on the
minimal known-good config FAILS LOUD instead of skipping. When unset (dev boxes,
unsupported cuDNN) it falls back to the normal skip guard so it never false-alarms.
The variable is intended for an operator to export on a known-good GPU CI leg once one exists. Note that today no ONNX Runtime pipeline definition exports it (there is no Hopper+ GPU CI leg), so it has no effect in this project's CI and only matters for manual/local runs where a developer sets it explicitly — don't describe it in test docstrings as an enforcement that CI already applies.
def require_cudnn_sdpa():
return os.environ.get("ORT_TEST_REQUIRE_CUDNN_SDPA") == "1"
# in the test:
enforce = require_cudnn_sdpa()
if not enforce and not cudnn_decode_supported(head_size): # illustrative: your suite's own support predicate
self.skipTest("cuDNN SDPA decode tier not dispatched; set ORT_TEST_REQUIRE_CUDNN_SDPA=1 to enforce")
# then assert dispatch == CUDNN_FLASH_ATTENTION unconditionallyRun both ways to prove it works AND bites:
python -m pytest <file>.py -v # normal: skips where unsupported
ORT_TEST_REQUIRE_CUDNN_SDPA=1 python -m pytest <file>.py -v # enforced: fails if not cuDNNProve the teeth. A canary you never watched fail is not verified. Force MATH-only
by setting the CUDA provider's sdpa_kernel provider option to the MATH bitmask
(16) — a monkeypatch of the C++ selector is not reachable from Python — under
ORT_TEST_REQUIRE_CUDNN_SDPA=1, and confirm it fails with, verbatim:
AssertionError: 'CUDNN_FLASH_ATTENTION' != 'MATH'A run that never demonstrates this failure has not proven the canary has teeth (grounding rule: negative/teeth evidence must actually be observed, not asserted).
WORKDIR=$(mktemp -d); cd "$WORKDIR" # neutral + private (see §1)
source /abs/repo/.venv/bin/activate
export LD_PRELOAD=/abs/cuda12.9/lib64/libcudart.so.12:/abs/cudnn9.8/lib/libcudnn.so.9
export LD_LIBRARY_PATH=/abs/cuda12.9/lib64:/abs/cudnn9.8/lib
export PYTHONPATH=/abs/repo/onnxruntime/test/python/transformers
F=/abs/repo/onnxruntime/test/python/transformers/<file>.py
python -m pytest "$F" -v # A: normal
ORT_TEST_REQUIRE_CUDNN_SDPA=1 python -m pytest "$F" -v # B: canary active, non-skippable
# C: teeth — force MATH under the env var, expect the AssertionError aboveCheck the passed count, not just the exit code. pytest -v exits 0 even if every
test skipped (no CUDA, or an unmet @skipUnless(ml_dtypes) guard) — the saved log then
looks like passing evidence but proves nothing. Require a non-zero passed count and
zero unexpected skips, and note pytest exit code 5 = "no tests collected" (usually a
wrong path or -k filter, not success). RUN B's canary only converts dispatch-related
skips into failures — it does not rescue collection or environment skips, so still
read the summary line.
Redirect to a log (... 2>&1 | tee "$WORKDIR/gpu_run.log") — the debug-info stdout and
pytest output are large, and a saved log is the evidence that the run happened and
dispatched to cuDNN. Write it inside $WORKDIR (the mktemp -d above), not a
predictable /tmp/gpu_run.log a co-tenant could pre-create as a symlink to clobber.
| Symptom | Root cause | Fix |
|---|---|---|
ModuleNotFoundError: onnxruntime.capi | repo-root ./onnxruntime/ source shadows the wheel | run pytest from a private mktemp -d (not repo root, not bare /tmp); abs path + PYTHONPATH |
routes to MATH / cuDNN load error | torch's bundled CUDA/cuDNN shadow ORT's | LD_PRELOAD system libcudart.so.12 + libcudnn.so.9, set LD_LIBRARY_PATH |
| test passes but path unproven | MATH fallback gives same numbers | observe SdpaKernel= via ORT_ENABLE_ATTENTION_KERNEL_DEBUG_INFO=1 |
| all tests skip green, regression hidden | dispatch-gated skip | ORT_TEST_REQUIRE_CUDNN_SDPA=1 makes assertions non-skippable |
© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .github/skills/ort-transformers-gpu-pytest of microsoft/onnxruntime.
Open the folder on GitHubat commit 8420709
ONNX Runtime GPU Transformers Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| ONNX Runtime GPU Transformers Tests this skillmicrosoft/onnxruntime | 22k | — | ~2.9k | Automated safety check: Pass | MIT | |
| Debug Failing GPUfacebookexperimental/triton | 201 | — | ~709 | Automated safety check: Pass | MIT | |
| Temporal Python Testingwshobson/agents | 40k | 11 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Squid Testing Pythoniusztinpaul/squid | 203 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Flaky Test DetectorArabelaTso/Skills-4-SE | 253 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Testing Pythonbenchflow-ai/skillsbench | 1.8k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 |
facebookexperimental/triton
Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.
wshobson/agents
Test Temporal workflows with pytest, time-skipping, and mocking strategies.
iusztinpaul/squid
Write and evaluate effective Python tests using pytest. An agent skill from iusztinpaul/squid.
ArabelaTso/Skills-4-SE
Identifies non-deterministic or unreliable tests through static code analysis and test result analysis.
benchflow-ai/skillsbench
Write and evaluate effective Python tests using pytest. An agent skill from benchflow-ai/skillsbench.
softspark/ai-toolkit
Testing strategy: pyramid, AAA, mocks/fakes/stubs, flaky tests, coverage.
microsoft/onnxruntime
Finds and fixes out-of-range output writes in ONNX Runtime operator shape-inference functions where a getNumOutputs guard admits too few outputs.
microsoft/onnxruntime
Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.
microsoft/onnxruntime
Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.
microsoft/onnxruntime
Triggers, re-runs and unblocks the CI checks on an ONNX Runtime pull request, after diagnosing whether a failure is transient or needs a code change.
microsoft/onnxruntime
Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component.
microsoft/onnxruntime
Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.
Categories
Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback. The skill specializes the general `ort-test` skill for the Python transformers tests under `onnxruntime/test/python/transformers/` on a machine that also has `torch` installed, and it lists gotchas that would each cost an hour to rediscover.capi'`.
ONNX Runtime GPU Transformers Tests fits situations like: running the ONNX Runtime transformers pytest suite against a GPU wheel; fixing a ModuleNotFoundError for onnxruntime.capi during tests; confirming a test really exercised the cuDNN SDPA path.
Run `npx skills add microsoft/onnxruntime --skill ort-transformers-gpu-pytest -a claude-code`. Or copy the skill folder (.github/skills/ort-transformers-gpu-pytest in microsoft/onnxruntime) into .claude/skills/ort-transformers-gpu-pytest in your project. Claude Code loads it when a task matches its description.
Run `npx skills add microsoft/onnxruntime --skill ort-transformers-gpu-pytest -a codex`. Or copy the skill folder (.github/skills/ort-transformers-gpu-pytest in microsoft/onnxruntime) into .agents/skills/ort-transformers-gpu-pytest in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/onnxruntime --skill ort-transformers-gpu-pytest -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ort-transformers-gpu-pytest, .gemini/skills/ort-transformers-gpu-pytest, .github/skills/ort-transformers-gpu-pytest and .opencode/skills/ort-transformers-gpu-pytest in your project.
Going by SKILL.md and its folder, ONNX Runtime GPU Transformers Tests needs the command-line tools its instructions call (python and pytest). Our summary lists: A GPU-built ONNX Runtime wheel installed in a virtual environment; pytest; A CUDA-capable GPU.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
ONNX Runtime GPU Transformers Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with ONNX Runtime GPU Transformers Tests: Debug Failing GPU (facebookexperimental/triton, 201 stars), Temporal Python Testing (wshobson/agents, 40k stars), Squid Testing Python (iusztinpaul/squid, 203 stars) and Flaky Test Detector (ArabelaTso/Skills-4-SE, 253 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/onnxruntime, which has 22,029 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 7, 2026.
Source: microsoft/onnxruntime on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.