Search

CUDA · microsoft/onnxruntime

5 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

microsoft/onnxruntime22k—~1.3kAutomated safety check: PassMITtoday
2

Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.

microsoft/onnxruntime22k—~1.4kAutomated safety check: PassMITtoday
3

Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component.

microsoft/onnxruntime22k—~1.6kAutomated safety check: PassMITtoday
4

Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

microsoft/onnxruntime22k—~2.9kAutomated safety check: PassMITtoday
5

Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing.

microsoft/onnxruntime22k—~6.5kAutomated safety check: PassMITtoday