Azure Pipelines Log Downloader
ansible/ansible
Downloads Azure Pipelines CI logs for an Ansible pull request or build so the agent can analyze test failures, after asking you first.
Triggers, re-runs and unblocks the CI checks on an ONNX Runtime pull request, after diagnosing whether a failure is transient or needs a code change.
$ npx skills add microsoft/onnxruntime --skill ort-ci -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/onnxruntime ort-ci --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/ort-ci .claude/skills/ort-ci && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ort-ci" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-ci into .claude/skills/ort-ci/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-ci", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-ciType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/onnxruntime --skill ort-ci -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/onnxruntime ort-ci --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.github/skills/ort-ci .agents/skills/ort-ci && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ort-ci" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-ci into .agents/skills/ort-ci/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-ci", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/onnxruntime --skill ort-ci -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/onnxruntime ort-ci --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.github/skills/ort-ci .cursor/skills/ort-ci && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ort-ci" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-ci into .cursor/skills/ort-ci/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-ci", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/onnxruntime.git --path .github/skills/ort-ci--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/onnxruntime --skill ort-ci -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/onnxruntime ort-ci --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.github/skills/ort-ci .gemini/skills/ort-ci && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ort-ci" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-ci into .gemini/skills/ort-ci/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-ci", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/onnxruntime ort-ciInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/onnxruntime --skill ort-ci -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .github/skills && cp -r skills-src/.github/skills/ort-ci .github/skills/ort-ci && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ort-ci" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-ci into .github/skills/ort-ci/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-ci", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/onnxruntime --skill ort-ci -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/onnxruntime ort-ci --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/onnxruntime.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.github/skills/ort-ci .opencode/skills/ort-ci && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ort-ci" agent skill from https://github.com/microsoft/onnxruntime/tree/main/.github/skills/ort-ci into .opencode/skills/ort-ci/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ort-ci", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ort-ciTriggers, re-runs and unblocks the CI checks on an ONNX Runtime pull request, after diagnosing whether a failure is transient or needs a code change.
Work happens on pull requests in `microsoft/onnxruntime`, where almost all CI runs as GitHub Actions workflows. Only the `Linux Android Emulator QNN CI Pipeline` still runs on Azure Pipelines, alongside the bot-driven `license/cla` status check. The agent starts by inspecting the PR's `statusCheckRollup` with `gh` so it does not queue duplicates, and tells GitHub Actions checks from Azure ones by their workflow name and details URL.
Failures are diagnosed before anything is re-run: most need a code change, and only genuinely transient problems such as network or disk errors deserve a manual re-run, and only while the PR head is unchanged. Checks that are already queued, running or successful are left alone, and after pushing a fix the agent stops because the push starts fresh CI. Triggers include stuck or missing checks, Python format failures, Doc Gen CI failures and `/azp run`.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a571b72. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ghgitpythonpipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gh, git and pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
ONNX Runtime CI Management loads about 4.1k tokens when it runs. Until then it costs about 136 tokens; SKILL.md has 1,531 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from microsoft/onnxruntime at commit a571b72, republished under its MIT licence (© microsoft). 1,531 words, ~4,113 tokens.
.claude/skills/ort-ci/SKILL.md (or your agent's skills folder).Workflows for triggering, re-running, and unblocking CI checks on an ONNX Runtime PR.
The repository is microsoft/onnxruntime. As of 2026-07, nearly all CI runs as GitHub
Actions workflows (~80+ checks per PR). Only one required check still runs on Azure
Pipelines — Linux Android Emulator QNN CI Pipeline (host aiinfra.visualstudio.com).
There is also the bot-driven license/cla status check.
Failures are not all the same. Before touching anything, diagnose each failure (see Triage: Diagnose Before Re-running): most failures need a code change and re-running them just fails again; only genuinely transient (network/disk) failures should be re-run via §1. §4 (Azure Pipelines) applies only to the single QNN pipeline.
Before doing anything, inspect current state so you do not queue duplicate runs.
If you make any change, commit and push it, then stop. A push updates the PR head SHA and automatically starts CI for the new commit. Do not manually re-run failures from the old SHA after pushing a fix; that only queues redundant runs against stale code. Manual re-runs are only for transient failures when the PR head has not changed.
GitHub Actions checks have a non-empty workflowName and a detailsUrl on github.com; the
Azure Pipelines check has an empty workflowName and a detailsUrl on aiinfra.visualstudio.com:
gh pr view <number> --repo microsoft/onnxruntime --json statusCheckRollup \
--jq '[.statusCheckRollup[] | {name, host:(.detailsUrl|split("/")[2])}]
| group_by(.host) | map({host:.[0].host, count:length})'# PR metadata + all checks grouped by state
gh pr view <number> --repo microsoft/onnxruntime \
--json number,title,url,state,isDraft,headRefName,headRefOid,baseRefName,statusCheckRollup
# Head / merge SHAs for external CI
gh api repos/microsoft/onnxruntime/pulls/<number> --jq '{head:.head.sha, merge:.merge_commit_sha}'Inspect statusCheckRollup and note, for each requested check, whether it is missing, queued,
in progress, failed, canceled, skipped, or already successful. Do not re-trigger a check
that is already queued/in_progress/SUCCESS unless the user explicitly asks.
Quickly list just the failed/pending checks:
gh pr view <number> --repo microsoft/onnxruntime --json statusCheckRollup \
--jq '.statusCheckRollup[]
| {name:(.name//.context), status:(.status//.state), conclusion:.conclusion}
| select(.conclusion!="SUCCESS" and .status!="SUCCESS")'Never blindly re-run failed CI. Most failures need a code change and will fail again identically on re-run. Only transient failures should be re-run. The process is: download the failed job's log, read the actual error, classify it, then fix or re-run case by case.
For a GitHub Actions check (the vast majority), get the run and read only the failed steps:
HEAD_SHA=$(gh api repos/microsoft/onnxruntime/pulls/<number> --jq .head.sha)
# Map failed checks to their workflow run IDs
gh run list --repo microsoft/onnxruntime --commit "$HEAD_SHA" --limit 100 \
--json databaseId,workflowName,status,conclusion,url \
--jq '.[] | select(.conclusion=="failure" or .conclusion=="cancelled")' # cspell:ignore cancelled -- literal GitHub API value
# Dump just the failed steps of a run (grep for the real error)
gh run view <run_id> --repo microsoft/onnxruntime --log-failed > /tmp/ci_<run_id>.log
grep -nE "error:|FAILED|warning:|Traceback|fatal error|No space left|Could not resolve|timed out" \
/tmp/ci_<run_id>.log | head -50For the Azure Pipelines QNN check, open its detailsUrl (a dev.azure.com /
aiinfra.visualstudio.com build page) and download the job log, or use the Azure DevOps
.../builds/<buildId>/timeline + log APIs (see the ci-failure-retrieval skill for the exact
requests).
| # | Failure class | How to recognize it in the log | Action — re-run or fix? |
|---|---|---|---|
| 1 | C/C++ warning-as-error | error: on a -Werror//WX line — e.g. implicit type-cast/narrowing (-Werror=conversion), unused variable/unused parameter (-Werror=unused-*), sign-compare, maybe-uninitialized | Fix code. Re-run will not help. Remove/[[maybe_unused]] the unused symbol, add an explicit static_cast<T>() / gsl::narrow_cast<T>() for the cast, or fix the real logic. Rebuild locally to confirm the warning is gone. |
| 2 | Test failure | [ FAILED ] Suite.Case (gtest) or FAILED test_*.py::... - AssertionError (pytest); often only on some EPs | Fix code/test. If a newly added op test fails only on EPs that don't support the op, restrict the test to supported EPs (e.g. gtest OpTester::Run(..., {kCpuExecutionProvider, kCudaExecutionProvider}) / excluded_provider_types, or skip via SetUp), or fix the kernel. Don't re-run unchanged. See the ort-test skill. |
| 3 | Transient / infra failure | Could not resolve host, Connection timed out, 429 Too Many Requests, No space left on device, package/download 5xx, agent lost, submodule clone timeout — with no compile/test error | Re-run (§1). This is the one class that a plain re-run fixes. If it recurs 2–3×, escalate — it may be a real infra/proxy issue, not noise. |
| 4 | Lint / Python format | Python format check fails; lintrunner reports diffs | Fix code with lintrunner -a, commit, push (§3). Re-run alone won't fix it. |
Rules of thumb:
error: or a [ FAILED ]/FAILED line means fix the code — re-running reruns
the same failing commit and fails identically.The repo ships a helper that re-runs only the GitHub Actions workflows whose latest run for the PR's current head commit failed/canceled — and skips any workflow that already has a newer run queued or in progress. Use it only after triage confirms the failures are transient (network/disk/agent) — see Triage. It is the safest way to retry those without piling on duplicates. Do not use this helper if you changed anything and pushed a new commit; the push already starts CI for the new head SHA.
Script: tools/scripts/rerun_failed_ci.sh
# Dry run first — shows what would be re-run, triggers nothing
./tools/scripts/rerun_failed_ci.sh <number> --dry-run
# Actually re-run the failed/canceled workflows for the PR's head commit
./tools/scripts/rerun_failed_ci.sh <number>
# Explicit repo (auto-detected from cwd when omitted)
./tools/scripts/rerun_failed_ci.sh <number> microsoft/onnxruntimeIt prefers gh run rerun <id> --failed (retry only failed jobs) and falls back to a full
rerun for fully canceled runs that have no discrete failed jobs. Requires an authenticated
gh. Always run --dry-run first and confirm the list looks right before the real run.
To re-run one specific workflow manually:
gh run list --repo microsoft/onnxruntime --commit <head_sha> --limit 100 \
--json databaseId,workflowName,status,conclusion,url
gh run rerun <run_id> --repo microsoft/onnxruntime --failed # only failed jobs
gh run rerun <run_id> --repo microsoft/onnxruntime # full rerunlicense/cla (CLA bot)The license/cla check is posted by Microsoft's CLA bot, independent of the CI pipelines.
When it is stuck as "Expected — Waiting for status to be reported", re-trigger only the bot —
no CI jobs are re-run — by posting this comment on the PR:
gh pr comment <number> --repo microsoft/onnxruntime \
--body "@microsoft-github-policy-service rerun"Then verify it flips to success:
gh pr view <number> --repo microsoft/onnxruntime --json statusCheckRollup \
--jq '.statusCheckRollup[] | select((.name//.context)=="license/cla")
| {status, conclusion}'Expect {"status":"COMPLETED","conclusion":"SUCCESS"}.
The required Python format check (job lint-python-format in
.github/workflows/lint.yml) runs
lintrunner --all-files and fails on any formatting/lint violation. Re-running it will not
help — you must fix the code, commit, and push. See
docs/Coding_Conventions_and_Standards.md
and the ort-lint skill.
# One-time setup (in an activated Python venv)
pip install -r requirements-lintrunner.txt
lintrunner init
# Auto-fix. Prefer changed files; use --all-files to match CI exactly.
lintrunner -a # changed files only
lintrunner -a --all-files # everything (what CI checks)
# Verify clean (no changes reported == pass)
lintrunner --all-filesThen commit and push the formatting fixes; the check re-runs automatically on the new commit:
git add -u && git commit -m "Fix lint" && git pushNotes:
lintrunner --all-files, so a local lintrunner -a on only changed files can miss a
pre-existing violation the CI reports. If the check still fails, run --all-files locally.lintrunner -a).As of 2026-07, the only ORT check still on Azure Pipelines is
Linux Android Emulator QNN CI Pipeline. Everything else is GitHub Actions (use §1). Trigger
it through the PR comment integration:
gh pr comment <number> --repo microsoft/onnxruntime \
--body "/azp run Linux Android Emulator QNN CI Pipeline"Then wait briefly and check for a reply from azure-pipelines[bot]:
# Note: through the GraphQL `comments` field (what `gh pr view --json comments`
# uses), the bot's author.login is `azure-pipelines` (no `[bot]` suffix), even
# though it surfaces as azure-pipelines[bot] in the UI and the REST API.
gh pr view <number> --repo microsoft/onnxruntime --json comments \
--jq '.comments[] | select(.author.login=="azure-pipelines")
| {createdAt, body}' | tailTrigger CI Pipelines section of the private gh-pr-management skill for the
dev.azure.com project/definition discovery and POST .../runs payload using
refName: refs/pull/<number>/merge and the PR merge_commit_sha)./azp run comment per attempt; do not spam repeated comments.The ONNX Runtime Windows GPU Doc Gen CI check (workflow
.github/workflows/windows_gpu_doc_gen.yml)
builds ORT and runs build.py --gen_doc validate. It fails when the generated operator docs
no longer match what's committed — typically after you add/modify an operator or its kernel
registrations but forget to regenerate docs/ContribOperators.md / docs/OperatorKernels.md.
Re-running will not help; you must update the docs.
The easiest fix is to download the regenerated docs the failed job already produced: on failure
the workflow uploads a single artifact named updated-docs that contains both
OperatorKernels.md and ContribOperators.md at its top level, so you can replace the committed
copies without building locally.
HEAD_SHA=$(gh api repos/microsoft/onnxruntime/pulls/<number> --jq .head.sha)
# Find the failed Doc Gen run
run_id=$(gh run list --repo microsoft/onnxruntime --commit "$HEAD_SHA" --limit 100 \
--json databaseId,workflowName,conclusion \
--jq '.[] | select(.workflowName|test("Doc Gen")) | select(.conclusion=="failure") | .databaseId' | head -1)
# Confirm the artifact is present (expect: updated-docs)
gh api repos/microsoft/onnxruntime/actions/runs/$run_id/artifacts --jq '.artifacts[].name'
# gh run download does not overwrite existing files. Remove both stale generated docs first,
# then extract their replacements (the artifact always holds both files).
rm docs/OperatorKernels.md docs/ContribOperators.md
gh run download "$run_id" --repo microsoft/onnxruntime -n updated-docs --dir docs/Then review, commit, and push — the check re-runs on the new commit:
git diff --stat docs/ContribOperators.md docs/OperatorKernels.md
git add docs/ContribOperators.md docs/OperatorKernels.md
git commit -m "Update operator docs" && git pushNotes:
updated-docs artifact always contains both OperatorKernels.md and
ContribOperators.md at its top level. Remove both committed files before downloading because
gh run download refuses to overwrite them; --dir docs/ then recreates both. Only the file(s)
whose generated content changed will show up in git diff after extraction.python tools/ci_build/build.py --config Release --build_dir build/Linux --gen_doc after a build, then commit the regenerated files. Downloading the
artifact is faster since it avoids a full build.| Check / job name | Owner | How to unblock |
|---|---|---|
license/cla | CLA bot | Comment @microsoft-github-policy-service rerun (§2) |
Python format (lint-python-format) | GitHub Actions | Fix with lintrunner -a, commit, push (§3) — rerun alone won't fix |
ONNX Runtime Windows GPU Doc Gen CI | GitHub Actions | Download the updated-docs artifact into docs/, commit, push (§5) — rerun alone won't fix |
Optional Lint, Optional Lint C++ | GitHub Actions | Non-required reviewdog checks; fix warnings or ignore |
Most CI (Linux/Windows/Mac/CUDA/TensorRT/WebGPU/Web/Android/iOS, windows_x64_*, Builds, PR Checks) | GitHub Actions | rerun_failed_ci.sh <number> (§1) |
Linux Android Emulator QNN CI Pipeline | Azure DevOps | /azp run Linux Android Emulator QNN CI Pipeline (§4), then API fallback |
ESRP), deployment, or
package-upload pipelines unless the user explicitly asks for that class of pipeline. Treat
names containing release, publish, official, nightly, sign, ESRP, production, or
deploy as high-risk and confirm first.--dry-run the rerun script before the real run.refs/pull/<number>/merge) over branch refs for external CI so the run
validates the merge result.statusCheckRollup.gh pr view <number> --repo microsoft/onnxruntime --json statusCheckRollup \
--jq '.statusCheckRollup | group_by(.status//.state)
| map({state:(.[0].status//.[0].state), count:length})'Confirm the previously stuck/failed check moved to queued/in_progress (or SUCCESS for the
CLA bot), and that no duplicate runs were created for the same head SHA.
© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .github/skills/ort-ci of microsoft/onnxruntime.
Open the folder on GitHubat commit a571b72
ONNX Runtime CI Management next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| ONNX Runtime CI Management this skillmicrosoft/onnxruntime | 22k | — | ~4.1k | Automated safety check: Pass | MIT | |
| Azure Pipelines Log Downloaderansible/ansible | 71k | — | ~825 | Automated safety check: Pass | GPL-3.0 | |
| CI Watchdoglatitude-dev/latitude-llm | 4.7k | — | ~1.6k | Automated safety check: Pass | MIT | |
| Migrate To TeamcityJetBrains/teamcity-cli | 125 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Megatron-LM CI Failure TriageNVIDIA/Megatron-LM | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| CI/CD Failure Troubleshootingruby-git/ruby-git | 1.8k | — | ~1.9k | Automated safety check: Pass | MIT |
ansible/ansible
Downloads Azure Pipelines CI logs for an Ansible pull request or build so the agent can analyze test failures, after asking you first.
latitude-dev/latitude-llm
Continuously monitor GitHub PR CI checks and automatically fix failures until all checks pass.
JetBrains/teamcity-cli
Migrating CI/CD pipelines to TeamCity. An agent skill from JetBrains/teamcity-cli.
NVIDIA/Megatron-LM
Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.
ruby-git/ruby-git
Diagnoses and fixes failing GitHub Actions runs by identifying the failure, fetching only the relevant logs, finding the root cause and reproducing it locally.
aiblueprinthq/ai-blueprint
Set up or normalize one project Verify command and matching GitHub Actions checks while preserving existing CI, with an optional local pre-push hook.
microsoft/onnxruntime
Finds and fixes out-of-range output writes in ONNX Runtime operator shape-inference functions where a getNumOutputs guard admits too few outputs.
microsoft/onnxruntime
Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.
microsoft/onnxruntime
Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.
microsoft/onnxruntime
Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component.
microsoft/onnxruntime
Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.
microsoft/onnxruntime
Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.
Works with
Categories
Triggers, re-runs and unblocks the CI checks on an ONNX Runtime pull request, after diagnosing whether a failure is transient or needs a code change. Work happens on pull requests in `microsoft/onnxruntime`, where almost all CI runs as GitHub Actions workflows. Only the `Linux Android Emulator QNN CI Pipeline` still runs on Azure Pipelines, alongside the bot-driven `license/cla` status check.
ONNX Runtime CI Management fits situations like: A required check on an ONNX Runtime PR is stuck, missing or failed; deciding whether a failing check is transient or needs a code change; fixing a failing Python format or operator-docs check; re-running the Azure Pipelines QNN check.
Run `npx skills add microsoft/onnxruntime --skill ort-ci -a claude-code`. Or copy the skill folder (.github/skills/ort-ci in microsoft/onnxruntime) into .claude/skills/ort-ci in your project. Claude Code loads it when a task matches its description.
Run `npx skills add microsoft/onnxruntime --skill ort-ci -a codex`. Or copy the skill folder (.github/skills/ort-ci in microsoft/onnxruntime) into .agents/skills/ort-ci in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/onnxruntime --skill ort-ci -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ort-ci, .gemini/skills/ort-ci, .github/skills/ort-ci and .opencode/skills/ort-ci in your project.
Going by SKILL.md and its folder, ONNX Runtime CI Management needs the command-line tools its instructions call (gh, git, python and pip). Our summary lists: GitHub CLI `gh`, signed in with access to the pull request.
SKILL.md contains no URLs. Its commands use gh, git and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
ONNX Runtime CI Management is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with ONNX Runtime CI Management: Azure Pipelines Log Downloader (ansible/ansible, 71k stars), CI Watchdog (latitude-dev/latitude-llm, 4.7k stars), Migrate To Teamcity (JetBrains/teamcity-cli, 125 stars) and Megatron-LM CI Failure Triage (NVIDIA/Megatron-LM, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/onnxruntime, which has 22,035 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 8, 2026.
Source: microsoft/onnxruntime on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.