Triagebot Action Bug Triage
withastro/astro
Takes a bug report for the triagebot-action GitHub Action through reproduction, root-cause diagnosis, an intended-behavior check and a fix attempt.
A skill your agent uses when asked to fix a GitHub issue end-to-end, run the agent pipeline on an issue, or process an agent:active / batch tracking issue.
$ npx skills add intel/torch-xpu-ops --skill issue-handler -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install intel/torch-xpu-ops issue-handler --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/issue-handler .claude/skills/issue-handler && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "issue-handler" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/issue-handler into .claude/skills/issue-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "issue-handler", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/issue-handlerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add intel/torch-xpu-ops --skill issue-handler -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install intel/torch-xpu-ops issue-handler --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/issue-handler .agents/skills/issue-handler && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "issue-handler" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/issue-handler into .agents/skills/issue-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "issue-handler", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add intel/torch-xpu-ops --skill issue-handler -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install intel/torch-xpu-ops issue-handler --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/issue-handler .cursor/skills/issue-handler && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "issue-handler" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/issue-handler into .cursor/skills/issue-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "issue-handler", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/intel/torch-xpu-ops.git --path .claude/skills/issue-handler--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add intel/torch-xpu-ops --skill issue-handler -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install intel/torch-xpu-ops issue-handler --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/issue-handler .gemini/skills/issue-handler && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "issue-handler" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/issue-handler into .gemini/skills/issue-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "issue-handler", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install intel/torch-xpu-ops issue-handlerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add intel/torch-xpu-ops --skill issue-handler -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/issue-handler .github/skills/issue-handler && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "issue-handler" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/issue-handler into .github/skills/issue-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "issue-handler", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add intel/torch-xpu-ops --skill issue-handler -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install intel/torch-xpu-ops issue-handler --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/issue-handler .opencode/skills/issue-handler && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "issue-handler" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/issue-handler into .opencode/skills/issue-handler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "issue-handler", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
issue-handlerA skill your agent uses when asked to fix a GitHub issue end-to-end, run the agent pipeline on an issue, or process an agent:active / batch tracking issue.
Issue Handler is an agent skill from intel/torch-xpu-ops, published by the product's own GitHub organization. Use when asked to fix a GitHub issue end-to-end, run the agent pipeline on an issue, or process an agent:active / batch tracking issue. Orchestrates the full pipeline: issue-triage → fix-reproduce → fix-root-cause → fix-implement → fix-verify → report. Handles single-bug issues and batch-bug issues (a parent tracking multiple sub-bugs — skip-list or heterogeneous — fanned out one fix branch per sub-bug). On re-run, prioritizes new human feedback and skips work that is still valid.
Its SKILL.md is about 8.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/execution-modes.md`).
It sits in Development, covering Issue triage, Root cause analysis and End-to-end testing. It works with GitHub and PyTorch. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0187b3b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitghpip3pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
download.pytorch.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Issue Handler loads about 8.8k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 128 tokens; SKILL.md has 3,944 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from intel/torch-xpu-ops at commit 0187b3b, republished under its Apache-2.0 licence (© intel). 3,944 words, ~8,822 tokens.
.claude/skills/issue-handler/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.This is the high-level scenario skill for handling a single GitHub issue. It does not do the detailed work itself; it sequences the leaf skills into one iterative pipeline and reports the result. Each stage's mechanics live in its own skill — read and follow that skill when you reach its stage.
Every agent-produced diff is a proposal. This skill commits the
staged fix onto a dedicated local branch agent/fix-issue-<N> (single
bug) or agent/fix-issue-<N>-<seq>-<slug> (per batch sub-item) and
records it in a fix_result*.json. When the bug is owned by an upstream
issue — a mirrored pytorch/pytorch#<M> DISABLED test, or a command
pointing at an upstream issue — the branch is named after that
number instead: agent/fix-pytorch-issue-<M>[-<slug>] (see
Branch naming). It never
pushes, tags, or opens a PR — the invoking workflow reads each
fix_result*.json, exports base_sha..branch as a patch, and a human
applies it. The leaves themselves only stage (they never commit).
Single-bug path (default for issue_type=single-bug):
triage → reproduce → root-cause → implement → verify → report
↑ |
└─── loop up to 3 times ───┘Batch path (issue_type=batch-bug, either batch_kind) — fan out the
single-bug pipeline per sub-item, each on its own fix branch:
triage → [preflight: install nightly wheel once]
→ for each sub-item:
reset + reproduce
├─ NOT_REPRODUCED → heterogeneous: ALREADY_FIXED · skip-list: STALE_SKIP (+follow-up)
└─ REPRODUCED → branch per [Branch naming](#branch-naming)
→ root-cause → implement → verify
(any failure marks the sub-item and continues the batch)
→ fan-out reportskip-list vs heterogeneous differ only in how a NOT_REPRODUCED
sub-item is labeled; the loop is identical.
| Stage | Leaf skill | Purpose |
|---|---|---|
| 1. Triage | issue-triage | Text-only classification: single-bug / batch-bug (+ batch_kind) / nonbug, scope, runtime_dependencies, preliminary verdict |
| 2. Reproduce | fix-reproduce | Verify the failure still reproduces against the nightly wheel (stage=nightly) |
| 3. Root cause | fix-root-cause | Deep source analysis, target_repo, domain, IMPLEMENTING/NEEDS_HUMAN |
| 4. Implement | fix-implement | Edit code, stage the diff (never commit) |
| 5. Verify | fix-verify | Run the refined command against source build, PASSED/FAILED/CANNOT_VERIFY |
| 6. Report | this skill | Summarize outcome to the user (or into the issue in pipeline mode) |
intel/torch-xpu-ops or pytorch/pytorch (URL,
number, or raw body).pytorch_dir — path to a local pytorch checkout, resolved as
described in fix-reproduce Prepare. If absent, this skill lets
fix-reproduce / fix-root-cause clone it into
$XPU_OPS_ROOT/agent_space_xpu/pytorch/.The pipeline runs in one of two modes — interactive (default) or pipeline — which changes how every stage reports results and whether it writes to the GitHub issue. Decide the mode at the start and pass it to every leaf. See references/execution-modes.md for the full contract.
agent:status marker, update stage labels, let leaf
skills leave their <!-- agent:<name> --> comments, and stop when
the pipeline settles on a terminal verdict.One rule for both paths: the branch is named after the issue that owns the bug, so a human reading a branch or a patch directory knows which issue a PR closes.
| The bug is filed as | Branch |
|---|---|
an intel/torch-xpu-ops issue #N | agent/fix-issue-<N> (batch sub-item: agent/fix-issue-<N>-<seq>-<slug>) |
an upstream pytorch/pytorch#<M> (mirrored DISABLED test, or a command naming an upstream issue) | agent/fix-pytorch-issue-<M> (batch sub-item: agent/fix-pytorch-issue-<M>-<slug>) |
The upstream form applies whenever the sub-item's own identity is an
upstream issue — the common case on the CI DISABLED tracking issue,
where #N is only a queue and #M is the bug. #M is already unique
and stable, so the upstream form carries no seq. When one fix covers
several sub-items, the branch is named after the sub-item that was fixed;
the rest are recorded in covers.
M is the upstream issue the sub-item resolves to: the
pytorch/pytorch#<M> reference on its checklist line, its linked child
reference, or the upstream issue named in the invoking comment. If a
sub-item has no upstream number, use the agent/fix-issue-<N>-… form.
Pipeline mode only; skip in interactive mode (a human is already
driving, so just run the full pipeline). An agent:active issue is
frequently re-triggered — a maintainer leaves feedback, or the bot is
re-invoked with no new information. This gate decides, before spending
a build, whether this is a fresh run, a human-feedback re-run
(highest priority), or a bare re-run (skip work that already ran and
is still valid).
Find the most recent agent comment and its timestamp:
last_agent_ts=$(gh issue view "$N" --repo "$OWNER/$REPO" --json comments \
--jq '[.comments[] | select(.body | test("<!-- agent:"))] | last | .createdAt // ""')Empty last_agent_ts → first run. Skip the rest of Stage 0 and go
to Stage 1 normally.
A comment is human feedback when it is authored by a non-bot account
and created after last_agent_ts. (The bot account is the one that
authored the <!-- agent:* --> comments; exclude it by login.)
new_human=$(gh issue view "$N" --repo "$OWNER/$REPO" --json comments \
--jq --arg ts "$last_agent_ts" --arg bot "$BOT_LOGIN" \
'[.comments[] | select(.author.login != $bot and .createdAt > $ts)] | length')new_human > 0 → human-feedback re-run. Human feedback is the
highest priority signal. Read every such comment verbatim and
prepend it to the failure description you hand to Stage 3
(fix-root-cause takes a free-form failure description; a leading
"Maintainer feedback since last run: ..." block steers the
re-analysis without any new leaf input). Run the full pipeline
from Stage 1; do not take any of the skip fast-paths below. A human
saying "still wrong" or "change X" overrides any cached verdict.new_human == 0 → bare re-run. Continue to Step 0.3.No human pointed anything out since the last agent comment, so re-doing the whole pipeline would just repeat identical work. Re-run only the cheap front of the pipeline and compare against last time:
refined_command + verdict are recoverable from the last
<!-- agent:root-cause --> comment's analyzed_sha context, or
re-derived by reading the prior <!-- agent:reproduce --> /
sweep comment. "Identical" means same verdict and same
refined_command.fix-root-cause's own
<!-- agent:root-cause --> analyzed_sha fast-path: if
target_repo HEAD sha equals the recorded analyzed_sha, that
leaf re-emits the prior verdict verbatim and this orchestrator
stops (the earlier outcome — fix already staged, or
NEEDS_HUMAN — still stands; there is nothing new to do). If the
sha moved, fix-root-cause re-analyzes and the pipeline continues
from Stage 3 as usual.This never skips Stage 1/Stage 2 — they are cheap (text + nightly wheel) and are the only way to notice the failure went away. It only avoids the expensive Stage 3-5 rebuild+fix when both the observable failure and the analyzed code are unchanged.
issue-triage)Call issue-triage on the issue body + comments. It emits
issue_type (single-bug / batch-bug / nonbug), batch_kind
(skip-list / heterogeneous / null), reproduction_missing
(yes / no), scope, runtime_dependencies, and a preliminary
handling (agent-fixable / needs-human).
Branch on issue_type (triage already made the batch-vs-nonbug call;
no re-detection here):
nonbug → stop the fix pipeline; skip to Stage 6 Report with
SKIPPED(reason=nonbug).single-bug with reproduction_missing=yes → stop; Stage 6 Report
with NEEDS_HUMAN(reason=reproduction_missing). issue-triage's
own comment already asks the reporter for a reproducer.single-bug with reproduction_missing=no → continue to Stage 2.batch-bug → Stage 1u, passing through batch_kind (the loop
uses it only to label a NOT_REPRODUCED sub-item).fix-reproduce)Only for the single-bug path (issue_type=single-bug). Call
fix-reproduce with:
reproducer_command — extracted by issue-triage from the issue
body.stage=nightly — reproduce against the nightly wheel.ci_repo — inferred from repo (torch-xpu-ops for issues on
intel/torch-xpu-ops, pytorch for pytorch/pytorch), or the value
the bot passes explicitly.Branch on its verdict:
REPRODUCED → continue to Stage 3. Record the refined_command
and base for downstream stages.NOT_REPRODUCED → the issue is stale; Stage 6 Report with
SKIPPED(reason=no_longer_reproduces).NO_REPRODUCER → Stage 6 Report with
NEEDS_HUMAN(reason=no_reproducer).CANNOT_VERIFY → Stage 6 Report with
NEEDS_HUMAN(reason=cannot_verify) and the blocker field.One loop for every parent issue that tracks multiple children. The parent lists several sub-items; run the single-bug pipeline on each independently, each on its own fix branch, and report every outcome back on the parent. "Fix what you can" — a sub-item that can't be fixed is marked and the batch continues.
Entered for issue_type=batch-bug. issue-triage already set
batch_kind; the two kinds share the entire loop and differ only in how
a NOT_REPRODUCED sub-item is labeled (see step 2 below):
heterogeneous — the parent body lists distinct sub-bugs. A
NOT_REPRODUCED entry means that bug is ALREADY_FIXED (no longer
reproduces); nothing to do.skip-list — a Bug Skip issue listing homogeneous already-skipped
tests. A NOT_REPRODUCED entry means the test now passes, so its skip
decorator is now stale and should be removed (follow-up).No child GitHub issues are created. Each sub-item is a checklist entry on the parent; its fix lives on a dedicated branch so a human can open one PR per fixed sub-item.
Parse the parent body into a list of sub-items. Each is either:
Class::method to a node id, fix-reproduce's Prepare step resolves
the file), orowner/repo#N — fetch that issue and use
its body as the sub-item's failure description. Only follow
intel/torch-xpu-ops and pytorch/pytorch references; ignore any
other repo (untrusted, per fix-root-cause).A sub-item may also arrive as a plain node id from the caller (a nightly
report, an email, a log excerpt): normalize Class::method to a node id,
drop blank lines, headers and comments, and collapse exact duplicates.
Split infra failures out before the loop. A runner crash, a docker
pull timeout, a full disk or a missing device is not a test failure and
will not reproduce. Record these as NEEDS_HUMAN(infra) and report
them; they never enter the loop.
Give each sub-item three identifiers used for branch naming:
seq — 1-based position in the body checklist order. Readable and
maps a branch back to its body line. Not the stable identity (editing
or reordering the body changes it).slug — the stable identity, derived from the sub-item's test node
id: take the leaf method name, strip device suffixes (_cpu / _xpu /
_cuda / _meta) and any dtype suffix, lowercase, keep [a-z0-9._-],
truncate to 40 chars. If two sub-items produce the same slug, append
-2, -3, … in body order so every slug is unique within the issue.upstream — the pytorch/pytorch#<M> number this sub-item is
filed as, from its checklist line or its linked child reference; empty
when there is none.Name the branch per Branch naming: an upstream
sub-item gets agent/fix-pytorch-issue-<M>-<slug>, otherwise
agent/fix-issue-<N>-<seq>-<slug> (N = parent issue number). On a
re-run, match a sub-item to its prior branch by slug — look for an
existing agent/fix-*-<slug> branch; if found, that is the same
sub-item (rename it if its seq moved, never create a duplicate). Skip
headers, prose, and empty lines.
When the list is long (skip-list issues routinely have dozens),
front-load the one real wheel install so the per-entry
fix-reproduce(stage=nightly) calls each find the env already current
and return fast. fix-reproduce always issues pip install --pre --upgrade (it refuses to trust a stale wheel); running it once here
does the real work. There is no skip_wheel_install flag — this is
purely an ordering optimization:
pip3 install --pre --upgrade torch torchvision torchaudio \
--index-url https://download.pytorch.org/whl/nightly/xpu
python -c "import torch; print('nightly:', torch.__version__)"Entries in a batch are often the same bug in the same class. Fixing each one separately runs root-cause and a source build (1-3 h) per entry and produces the same patch several times.
Do not guess which entries share a fix:
REPRODUCED entries so likely relatives sit together (same
test class or file, same error). Ordering only; it decides nothing.fix-verify
already produces a before/after table; extend it to the group.covers list and are not fixed
again. Entries that still fail take their own turn from step 2.One fix gets one branch and one fix_result.json, naming in covers
every entry it also fixed.
Capture the two base SHAs once, then for each sub-item reset both checkouts per the shared reset-between-entries recipe — a prior sub-item's staged diff must not bleed into the next.
For each sub-item:
Reset both checkouts to the base SHAs (shared recipe).
Reproduce (fix-reproduce). It detaches the pytorch tree to its
base (origin/main or the ci_commit fallback) and returns that
base plus a refined_command.
stage follows the issue, like allow_skip:
pytorch-ci-failure, or a mirrored DISABLED test) →
stage=auto. Many CI failures only reproduce on a source build.
stage=nightly returns NOT_REPRODUCED as soon as the nightly
wheel passes and does not fall through, which records unfixed work
as already fixed.stage=nightly.Branch on the verdict:
REPRODUCED → continue to step 3; keep its base and
refined_command.NOT_REPRODUCED → nothing to fix; record it and go to the next
sub-item. The label depends on batch_kind:heterogeneous → ALREADY_FIXED: the reported bug no longer
reproduces on latest nightly (upstream fixed it, or it was flaky).
No action needed.skip-list → STALE_SKIP: the test now passes, so its skip
decorator is obsolete. Record a follow-up to remove the
decorator — the orchestrator does not delete it here (see
"Stale skips" below).NO_REPRODUCER → INVALID_ENTRY (renamed/removed test, or
malformed). Record, continue.CANNOT_VERIFY → UNVERIFIED (environmental). Record, continue.Root-cause, then branch. Run Stage 3 (fix-root-cause)
first — it returns target_repo, which decides which checkout the
fix (and thus the branch) lives in:
target_repo=pytorch → target_repo_dir = pytorch_dir, base is
the reproduce base on the pytorch tree.target_repo=torch-xpu-ops →
target_repo_dir = pytorch_dir/third_party/torch-xpu-ops, base is
xpu_ops_base (the submodule's pinned commit from the shared
recipe). fix-implement / fix-verify handle the xpu.txt pin
rewrite internally; the orchestrator only creates the branch.Create the isolated fix branch on target_repo_dir off its base so
the diff is pushable on its own (Branch naming):
# upstream sub-item: branch="agent/fix-pytorch-issue-${M}-${slug}"
branch="agent/fix-issue-${N}-${seq}-${slug}"
git -C "$target_repo_dir" checkout -B "$branch" "$base"Then run Stage 4 → 5 (same contract and 3-attempt bound as the
single-bug path). fix-implement expects exactly this: a fresh
branch the orchestrator just created, clean worktree.
On any leaf NEEDS_HUMAN / CANNOT_VERIFY / FAILED (or
attempts exhausted): mark this sub-item blocked with the reason,
record it, and continue to the next sub-item — never abort the
whole batch on one hard sub-item.
On fix-verify PASSED: do NOT push or open a PR. Update the
sub-item's fix_result-${slug}.json in place (written at
verdict=PENDING_VERIFY when the branch was committed) to
verdict=PASSED with fix_repo_dir, branch, base_sha (the
sub-item's base from step 3), changed_files,
so the workflow can export base_sha..branch per sub-item. As in the
single-bug path, commit the branch and write the PENDING_VERIFY
record as soon as fix-implement returns READY, not only on PASSED,
so a crash mid-sub-item leaves a recoverable record instead of an
orphan branch.
batch_kind=skip-list only)A STALE_SKIP sub-item's skip decorator can be removed, but this
orchestrator does not delete it. Deleting a skip is itself a code
change that needs its own verify + PR; folding it into this sweep would
mix "the skip is obsolete" with "and here's the removal diff" and
obscure the batch outcome. Instead, surface every STALE_SKIP in the
report as an explicit follow-up (candidate for a human, or a separately
invoked fix-implement run to remove the decorator). The report is the
hand-off; nothing is auto-removed.
Post one summary comment on the parent (or surface to the user in interactive mode), and mirror it into the parent's checklist:
<!-- agent:batch-fanout -->
## Batch fan-out results
Base: <torch nightly version or base sha>
| Sub-item | Outcome | Branch / Reason |
|---|---|---|
| test_bar_xpu_float32 | FIXED | agent/fix-issue-4321-1-test_bar |
| test_baz | NEEDS_HUMAN | cross_repo_coordinated |
| test_qux | NEEDS_HUMAN | attempts_exhausted |
| test_new | ALREADY_FIXED | no longer reproduces on latest nightly |
| test_old | STALE_SKIP | follow-up: remove skip decorator |
| test_dup | COVERED | by test_bar's fix, agent/fix-issue-4321-1-test_bar |
| test_hard | SKIPPED | skip added, tracking issue intel/torch-xpu-ops#1234 |
| test_gone | INVALID_ENTRY | does not collect |
- **FIXED:** N sub-items — one branch each, ready for a human to open
a PR.
- **NEEDS_HUMAN:** M sub-items — see per-sub-item reason.
- **ALREADY_FIXED:** J sub-items (heterogeneous only) — no longer
reproduce; the reported bug is resolved, no action needed.
- **STALE_SKIP:** K sub-items (skip-list only) — the test now passes;
follow up to remove the obsolete skip decorator.
- **INVALID_ENTRY / UNVERIFIED:** P sub-items — malformed/renamed, or
environmental during reproduce.
*Automated by issue-handler.*Omit category rows that have no members (a heterogeneous batch has
ALREADY_FIXED but never STALE_SKIP, and vice versa). The
<!-- agent:batch-fanout --> marker lets a
re-run locate and update this same comment in place. On a re-run
(Stage 0), only re-process sub-items that are not already FIXED on a
live branch, unless human feedback (Stage 0.2) reopens a specific one.
After the loop, go to Stage 6 Report with the aggregate outcome
(IMPLEMENTING(fix_verified) if any sub-item was fixed;
NEEDS_HUMAN only if every actionable sub-item needed a human —
STALE_SKIP follow-ups do not by themselves force NEEDS_HUMAN).
Written on both paths: the single-bug Stage 4/5 writes
fix_result.json, and a batch fan-out writes batch_summary.json plus
one fix_result-<slug>.json per fixed sub-item. The
<!-- agent:batch-fanout --> comment is for humans; these files let the
invoking bot workflow export patches and drive re-verification without
re-parsing a comment.
They live under $AGENT_SPACE (the gitignored scratch dir). In pipeline
mode the bot workflow always sets it. In interactive mode it is usually
unset, which would resolve to /fix_result.json — so fall back to
<repo>/agent_space_xpu/ (the same gitignored dir AGENTS.md defines) and
report the path you used:
agent_space="${AGENT_SPACE:-$(git -C "$target_repo_dir" rev-parse --show-toplevel)/agent_space_xpu}"
mkdir -p "$agent_space"batch_summary.json — one file listing every sub-item:
{
"issue": 4321,
"kind": "batch-bug",
"batch_kind": "heterogeneous",
"sub_items": [
{ "seq": 1, "slug": "test_bar", "outcome": "FIXED",
"branch": "agent/fix-issue-4321-1-test_bar",
"target_repo": "torch-xpu-ops",
"fix_result": "fix_result-test_bar.json",
"summary": "one-line what/why" },
{ "seq": 2, "slug": "test_baz", "outcome": "NEEDS_HUMAN",
"branch": null, "reason": "cross_repo_coordinated" },
{ "seq": 3, "slug": "test_new", "outcome": "ALREADY_FIXED",
"branch": null, "reason": "no longer reproduces on latest nightly" }
]
}fix_result-<slug>.json — for each FIXED sub-item, the same
schema the single-bug Stage 5 writes as fix_result.json. It MUST
include the keys the workflow's schema gate and patch-export read:
verdict — PASSED / FAILED / CANNOT_VERIFY / PENDING_VERIFY.
Written incrementally: Stage 4 writes PENDING_VERIFY when it commits
the branch, Stage 5 updates it to the terminal verdict. Only PASSED
is exported as a normal patch; a non-PASSED verdict that still names a
reachable branch + base_sha is salvaged as an unverified/ patch
so the committed work is not lost with the runner. Keep branch and
base_sha accurate even on a non-PASSED record.target_repo — torch-xpu-ops | pytorch,fix_repo_dir — absolute path of the git repo holding the fix commit,branch — per Branch naming,base_sha — the commit the fix branch was started from,changed_files — list of the files the fix touched,
plus needs_build, refined_command, notes. Suffixed by slug so each
sub-item can be exported / re-verified independently.The patch-export step iterates every fix_result*.json and emits one
patch series per unit from base_sha..branch (the branches are not
pushed); batch_summary.json enriches it (target_repo, summary line).
A unit that reports a verified fix but yields no patch fails the step.
fix-root-cause)Called on the single-bug path, and once per sub-item from Stage 1u.
Call fix-root-cause with the failure description and the
refined_command from Stage 2.
Branch on its verdict:
IMPLEMENTING(reason=ok) → record target_repo, domain,
analyzed_sha, root_cause, fix_strategy, then create the fix
branch before Stage 4 — target_repo is what decides which checkout
it lives in, and fix-implement expects a fresh branch with a clean
worktree ($base is Stage 2's reproduce base; the batch path already
did this in its own step 3, with its own branch name):
# upstream-filed bug: branch="agent/fix-pytorch-issue-${M}"
branch="agent/fix-issue-${N}"
git -C "$target_repo_dir" checkout -B "$branch" "$base"NEEDS_HUMAN → Stage 6 Report with the specific
reason (task_or_feature / feature_gap / hardware_specific /
cross_repo_coordinated / no_registered_domain / etc.). Each
reason maps to a different final agent:status value; see
execution-modes.md.
fix-implement)Call fix-implement with triage_result, pytorch_dir,
target_repo_dir (derived from target_repo), and allow_skip:
allow_skip follows the issue:
pytorch-ci-failure label, or
mirrors a DISABLED test from pytorch/pytorch. allow_skip=true. Fix
it in place where you can, in pytorch or here. Skip only when the fix
needs a dependency, information you do not have, or feature-sized
work; fix-implement then files a tracking issue for the real fix.allow_skip=false. Never add a skip decorator.If the caller states the flag, the caller wins.
Branch on the verdict:
READY(reason=ok) → commit the staged fix onto the branch created in
Stage 3, and write the incremental fix_result.json, then continue
to Stage 5.
Do NOT wait for Stage 5 to persist — if verify then crashes or the
context runs out, the committed branch would otherwise be an
orphan the Export step flags as a lost fix. Committing here keeps the
branch and the hand-off record in lock-step:
git -C "$target_repo_dir" commit -m "fix: <one-line summary> (#${N})"Then write fix_result.json (see "Machine-readable outputs" for the
path and its interactive-mode fallback) with the fields known so far
and verdict=PENDING_VERIFY: target_repo, fix_repo_dir, branch,
base_sha, changed_files, analyzed_sha, root_cause, notes.
The branch is local and dies with the runner, so this record is what
lets the workflow salvage the commit as an unverified/ patch if
verification never finishes — without it, a crash after this point
loses the work outright.
NEEDS_HUMAN → Stage 6 Report. The specific reason
(skip_outside_target_repo / skip_guard_rejected /
no_fix_possible / etc.) drives the final label.
fix-verify)Call fix-verify with refined_command (from Stage 2),
target_repo_dir, and changed_files (from Stage 4). fix-verify
unconditionally produces the FAIL->PASS before/after table and runs
spin fixlint on a passing result — no flags to pass.
Branch on the verdict:
PASSED(reason=ok) → the fix is verified. The branch and
fix_result.json already exist from Stage 4; update the record
in place — set verdict=PASSED and add needs_build, the
before/after result, and any final notes — then go to Stage 6 with
IMPLEMENTING(fix_verified). fix-verify leaves its spin fixlint
fixes staged: amend them into the Stage 4 commit
(git commit --amend --no-edit) so they reach the exported patch —
never a second commit. The workflow exports base_sha..branch as the
patch. Never push or open a PR.FAILED → loop back to Stage 4 with the failure output as
additional context (see "Iterative loop bounds"). On a retry the
branch already carries the previous attempt's commit, so amend it
(git commit --amend --no-edit after staging the new edits) rather
than creating a second commit or a new branch — base_sha..branch
must stay a single reviewable fix, and Stage 5's clean-worktree
assertion requires nothing left uncommitted. If attempts are
exhausted, update fix_result.json to verdict=FAILED with the last
failure in notes so the record explains the branch, then Stage 6
Report NEEDS_HUMAN(reason=attempts_exhausted).CANNOT_VERIFY → update fix_result.json to
verdict=CANNOT_VERIFY with the blocker in notes, then Stage 6
Report NEEDS_HUMAN(reason=<verify's reason>). Do not loop on
CANNOT_VERIFY — the environment problem will not fix itself.At the end, summarize the outcome. In interactive mode present
this to the user in plain language. In pipeline mode advance
the issue's agent:status to the terminal stage
(DONE / NEEDS_HUMAN / SKIPPED) and update the checklist per
execution-modes.md; a batch fan-out
run also posts the <!-- agent:batch-fanout --> summary.
Always include in the summary:
single-bug / batch (with batch_kind).IMPLEMENTING(fix_verified) / NEEDS_HUMAN(<reason>)
/ SKIPPED(<reason>).STALE_SKIP follow-ups when batch_kind=skip-list.If the outcome is IMPLEMENTING(fix_verified), the fix is committed on
its fix branch (Branch naming) and recorded in
fix_result.json. The invoking workflow reads that, exports
base_sha..branch as a patch
artifact, and a human applies it. Do not push or open the PR from this
skill.
batch-fanout and summary block templatesBoth are this skill's own blocks, so write them collapsed as shown —
there is no ## heading to strip, unlike the leaf blocks this skill
wraps.
Batch runs only:
<!-- agent:batch-fanout -->
<details>
<summary><b>Batch fan-out results</b></summary>
Base: `torch-xpu-ops@<short_sha>`, pytorch `origin/main@<short_sha>`
(reproduce base: nightly `<wheel version>`)
| Sub-item | Outcome | Branch / Reason |
|---|---|---|
| 4. `test_foo_xpu_float8_e4m3fn` | FIXED | `agent/fix-pytorch-issue-197334-test_foo` |
| 5. `test_foo_xpu_float8_e5m2` | COVERED | by sub-item 4's fix, same branch |
- **FIXED:** <n> sub-items — <what a human should do with the branches>
- **COVERED:** <n> sub-items — <why no separate branch>
- **NEEDS_HUMAN / STALE_SKIP / INVALID_ENTRY / UNVERIFIED:** none.
*Automated by issue-handler.*
</details>Every run. The sign-off is outside </details> — it is the comment's
footer, not part of the summary:
<!-- agent:summary -->
<details>
<summary><b>Summary</b></summary>
<2-4 sentences: the one bug behind the sub-items, the fix in one clause,
and what verification showed. No bullets.>
- **Branch:** `agent/fix-issue-<N>[-<seq>-<slug>]`, or
`agent/fix-pytorch-issue-<M>[-<slug>]` for an upstream-filed bug
- **Base:** `<full 40-char sha>`
- **Changed files:** `path/to/file.cpp`
- **Nothing was pushed.** The patch is exported as a workflow artifact for
a human to review and apply.
</details>
_Generated by [fix job](<run url>)._When the job ends, the fix workflow appends the <!-- agent:review -->
block itself — the verdict menu a reviewer copies into a new comment.
Do not write it yourself, and do not restate it here: the template
lives in bot.yml and a second copy only drifts from it.
The pipeline is not strictly linear. Loop when a later stage invalidates an earlier assumption:
FAILED → return to Stage 4 (refine the fix).NEEDS_HUMAN and the
orchestrator does not retry.Bound: maximum 3 fix attempts (Stage 4 → Stage 5 → Stage 4 …).
This matches the legacy pipeline's max_agent_attempts. When you
stop without success, report NEEDS_HUMAN(reason=attempts_exhausted)
with the last fix-verify failure output in reason_detail.
Do not loop on:
CANNOT_VERIFY at any stage (environment problem, not fix
problem).NEEDS_HUMAN from any leaf (contract: the leaf already decided
it needs a human).no_registered_domain (domain registry is a fixed set,
looping won't unstick it).Pipeline mode only. In interactive mode, do not touch the issue body/markers/labels unless the user asks — report to the user instead.
This orchestrator owns advancing the overall <!-- agent:status:X -->
marker through:
DISCOVERED → TRIAGING → REPRODUCING → TRIAGED → IMPLEMENTING →
VERIFYING → DONEwith terminal alternates NEEDS_HUMAN and SKIPPED. Stage-by-stage
mapping to labels is in
references/execution-modes.md; each
leaf skill owns its own <!-- agent:<name> --> comment/log slot,
this orchestrator owns the overall agent:status marker + the
Action Items checklist.
© intel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in .claude/skills/issue-handler of intel/torch-xpu-ops.
Open the folder on GitHubat commit 0187b3b
Issue Handler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Issue Handler this skillintel/torch-xpu-ops | 115 | — | ~8.8k | Automated safety check: Pass | Apache-2.0 | |
| Triagebot Action Bug Triagewithastro/astro | 63k | — | ~639 | Automated safety check: Pass | Custom licence | |
| Issue TracerZaxbyHub/opencode-swarm | 488 | — | ~4.4k | Automated safety check: Pass | MIT | |
| Issue Backlog Clusteringthedotmack/claude-mem | 97k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Triagearcee-ai/nac | 279 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| LazyCodex Bug Reportercode-yeongyu/oh-my-openagent | 70k | — | ~3k | Automated safety check: Pass | Custom licence |
withastro/astro
Takes a bug report for the triagebot-action GitHub Action through reproduction, root-cause diagnosis, an intended-behavior check and a fix attempt.
ZaxbyHub/opencode-swarm
Drives a bug report from validation and root-cause tracing through a critic-reviewed plan, an approved minimal fix and a PR-ready closure, never merging without recorded human approval.
thedotmack/claude-mem
Groups a large GitHub issue backlog by root cause into plan-master issues, redirects the child issues, and bundles one PR per cluster that closes them together.
arcee-ai/nac
Triage a GitHub repository's open issues by finding exact duplicates, rejecting evidenceably off-base requests, requesting concrete clarification, applying only existing labels, and opening a linked…
code-yeongyu/oh-my-openagent
Investigates a LazyCodex or Codex CLI defect, decides which GitHub repository owns it, and drafts an evidence-backed issue or pull request with repro steps.
warpdotdev/oz-for-oss
Triage a newly filed GitHub issue in this repository by analyzing the report, inspecting relevant code, estimating reproducibility, suggesting the likely root cause, and returning structured triage…
intel/torch-xpu-ops
Select the Intel GPU device to use when a system has multiple Intel GPU devices.
intel/torch-xpu-ops
Check PyTorch ciflow/xpu (xpu.yml) on the main branch, collect the failing XPU test cases from the most recent completed run(s), analyze the ROOT CAUSE of each failure with AI, and produce a list…
intel/torch-xpu-ops
Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.
intel/torch-xpu-ops
Review pull requests for XPU operator or backend code. An agent skill from intel/torch-xpu-ops.
intel/torch-xpu-ops
Guide users through creating Agent Skills for Claude Code. An agent skill from intel/torch-xpu-ops.
intel/torch-xpu-ops
Read the evidence a nightly UT run produced, decide which failures share a root cause and which are machine breakage rather than product bugs, and write one issue draft per root cause to drafts.json.
Categories
A skill your agent uses when asked to fix a GitHub issue end-to-end, run the agent pipeline on an issue, or process an agent:active / batch tracking issue. Issue Handler is an agent skill from intel/torch-xpu-ops, published by the product's own GitHub organization. Use when asked to fix a GitHub issue end-to-end, run the agent pipeline on an issue, or process an agent:active / batch tracking issue.
Issue Handler fits situations like: asked to fix a GitHub issue end-to-end; run the agent pipeline on an issue; process an agent:active / batch tracking issue.
Run `npx skills add intel/torch-xpu-ops --skill issue-handler -a claude-code`. Or copy the skill folder (.claude/skills/issue-handler in intel/torch-xpu-ops) into .claude/skills/issue-handler in your project. Claude Code loads it when a task matches its description.
Run `npx skills add intel/torch-xpu-ops --skill issue-handler -a codex`. Or copy the skill folder (.claude/skills/issue-handler in intel/torch-xpu-ops) into .agents/skills/issue-handler in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intel/torch-xpu-ops --skill issue-handler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/issue-handler, .gemini/skills/issue-handler, .github/skills/issue-handler and .opencode/skills/issue-handler in your project.
Going by SKILL.md and its folder, Issue Handler needs the command-line tools its instructions call (git, gh, pip3 and python).
SKILL.md names 1 domain. In commands or code: download.pytorch.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Issue Handler is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.8k tokens (SKILL.md is roughly 35k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Issue Handler: Triagebot Action Bug Triage (withastro/astro, 63k stars), Issue Tracer (ZaxbyHub/opencode-swarm, 488 stars), Issue Backlog Clustering (thedotmack/claude-mem, 97k stars) and Triage (arcee-ai/nac, 279 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
intel (a GitHub organization, an official publisher) maintains it in intel/torch-xpu-ops, which has 115 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 6, 2026.
Source: intel/torch-xpu-ops on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.