The Art of Debugging
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
A skill your agent uses when asked to implement a fix for a root-caused failure, apply a proposed patch, or produce a staged code change from triage output.
$ npx skills add intel/torch-xpu-ops --skill fix-implement -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install intel/torch-xpu-ops fix-implement --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/fix-implement .claude/skills/fix-implement && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "fix-implement" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/fix-implement into .claude/skills/fix-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-implement", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/fix-implementType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add intel/torch-xpu-ops --skill fix-implement -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install intel/torch-xpu-ops fix-implement --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/fix-implement .agents/skills/fix-implement && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "fix-implement" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/fix-implement into .agents/skills/fix-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-implement", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add intel/torch-xpu-ops --skill fix-implement -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install intel/torch-xpu-ops fix-implement --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/fix-implement .cursor/skills/fix-implement && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "fix-implement" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/fix-implement into .cursor/skills/fix-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-implement", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/intel/torch-xpu-ops.git --path .claude/skills/fix-implement--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add intel/torch-xpu-ops --skill fix-implement -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install intel/torch-xpu-ops fix-implement --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/fix-implement .gemini/skills/fix-implement && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "fix-implement" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/fix-implement into .gemini/skills/fix-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-implement", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install intel/torch-xpu-ops fix-implementInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add intel/torch-xpu-ops --skill fix-implement -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/fix-implement .github/skills/fix-implement && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "fix-implement" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/fix-implement into .github/skills/fix-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-implement", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add intel/torch-xpu-ops --skill fix-implement -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install intel/torch-xpu-ops fix-implement --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/torch-xpu-ops.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/fix-implement .opencode/skills/fix-implement && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "fix-implement" agent skill from https://github.com/intel/torch-xpu-ops/tree/main/.claude/skills/fix-implement into .opencode/skills/fix-implement/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fix-implement", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
fix-implementA skill your agent uses when asked to implement a fix for a root-caused failure, apply a proposed patch, or produce a staged code change from triage output.
Fix Implement is an agent skill from intel/torch-xpu-ops, published by the product's own GitHub organization. Use when asked to implement a fix for a root-caused failure, apply a proposed patch, or produce a staged code change from triage output. Takes fix-root-cause's output and edits code; leaves the change staged (uncommitted) for fix-verify to check. Does NOT run tests, does NOT commit, does NOT open PRs. Called by issue-handler, with allowskip set from the kind of issue being fixed.
Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Development, covering Root cause analysis. It works with PyTorch. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0187b3b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitghFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git and gh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Fix Implement loads about 4.2k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 1,862 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from intel/torch-xpu-ops at commit 0187b3b, republished under its Apache-2.0 licence (© intel). 1,862 words, ~4,154 tokens.
.claude/skills/fix-implement/SKILL.md (or your agent's skills folder).Takes triage output and makes the code change. Does not run tests (that
is fix-verify's job); staging-only, see HARD RULES.
triage_result — JSON output from fix-root-cause (root_cause,
fix_strategy, target_repo, domain, analyzed_sha).PYTORCH_DIR — path to local PyTorch checkout.target_repo_dir — path to the checkout that will be edited. Derived
from PYTORCH_DIR and triage_result.target_repo:target_repo == "pytorch" → target_repo_dir = PYTORCH_DIR.target_repo == "torch-xpu-ops" →
target_repo_dir = <PYTORCH_DIR>/third_party/torch-xpu-ops
(per AGENTS.md "Commit Pin & Development Override"; the caller
is expected to have cloned the working branch there ahead of
time).
All git and edit operations in this skill run against
target_repo_dir, never against PYTORCH_DIR when the two differ.allow_skip — controls skip decorator strategy:false (UT and everything else): never add skip decorators; must
unskip and really fix.true (CI break: pytorch-ci-failure, or a mirrored DISABLED
test): may add a skip with a tracking issue when
the fix needs dependencies, information you do not have, or
feature-sized work. Repair in place first; the skip is the fallback.
Stale skips must still be removed.basename $(git -C $target_repo_dir rev-parse --show-toplevel) # confirm which repo
git -C $target_repo_dir status # inspect current stateOn the first attempt (fresh branch just created by the orchestrator), the worktree is clean.
On a loop-back attempt (orchestrator returned here after
fix-verify FAILED or the reviewer requested changes), staged
changes from the previous attempt are still present by design. Do
NOT abort. Refine those changes on top; do not git reset or
git clean — the orchestrator would have done it if a fresh start
were intended.
Read triage_result carefully before touching any file. Understand:
analyzed_sha the triage was against — your edits go on top
of that baseIf the issue is not yet triaged, run fix-root-cause first.
../domain-knowledge/domain-<name>.md, per the domain
registry) for path conventions.See fix-root-cause Step 1 for domain routing. Common strategies:
atol/rtol values exactly.git log --oneline -20 -- <file>),
apply a fix aligned with upstream intent; document any divergence in
comments.allow_skip=false
and support is genuinely missing, report NEEDS_HUMAN — do not add
a skip. If allow_skip=true, follow the "Add a new skip" recipe in
"Skip operations" below.Skip decorators for a failing test live wherever that test lives: the
pytorch test tree for upstream tests, test/xpu/ inside torch-xpu-ops
for its own tests.
The skip must live inside target_repo_dir. This skill only ever
produces a single-repo diff, and the orchestrator commits (or reads)
only target_repo_dir. If the skip would have to be added to a file
outside target_repo_dir — e.g. target_repo == "torch-xpu-ops" but
the failing test is in pytorch's test/ tree — do NOT edit it: that
change would be left uncommitted and then wiped by the orchestrator's
reset between failures. Return NEEDS_HUMAN(reason=skip_outside_target_repo)
naming the file that would need the skip.
Removing a stale skip (a @skipIfXpu / xfail decorator whose
underlying failure this run is fixing): read the test file, delete
the decorator lines, save. Nothing else to do — the change goes
through the normal git add in Step 3.
Adding a new skip (allow_skip=true only). File a tracking issue
first so the skip has a follow-up owner, then edit:
# 1. Create the tracking issue and capture the URL.
issue_url=$(gh issue create \
--repo intel/torch-xpu-ops \
--title "[skip-added] <test_id> on XPU" \
--body "Auto-added by fix-implement (allow_skip=true).
Test: <test_id>
Original failure: <one-line failure summary from triage_result>
Root cause: <root_cause from triage_result>
Reason for skip: <why the actual fix requires human follow-up>
Base analyzed: <target_repo>@<short_sha from analyzed_sha>
" \
--label "agent-added,module: xpu" \
| tail -1)
# 2. Add the decorator with a comment citing $issue_url so a human
# can find the tracking issue by grepping the skip in place.
# Example additions (choose the one matching the test's shape):
#
# @skipIfXpu(f"see {issue_url}")
# def test_foo(self): ...
#
# # or in an OpInfo skips tuple:
# DecorateInfo(unittest.skip(f"see {issue_url}"), 'TestFoo', 'test_bar', ...)Emit the resulting issue_url as tracking_issue in the JSON output
(see Output section). skip_added becomes true. The issue label
agent-added lets a human filter for automated-triage tracking issues.
git -C $target_repo_dir add <your_files>
git -C $target_repo_dir diff --cached --stat # verify only intended files are stagedNever stage unrelated files. In particular, never stage
third_party/xpu.txt — it is a submodule pin managed by build
tooling, not part of any bug fix (see HARD RULES).
allow_skip=false)Skip this step entirely when allow_skip=true (the CI-break flow).
When allow_skip=false, spawn a fresh-context subagent via the Task
tool (subagent_type=general-purpose) to inspect the staged diff for
skip-shaped workarounds before returning to the orchestrator. The
implementer must not review its own diff for this specific bias — the
gatekeeper here is a separate agent with no memory of the reasoning
that produced the diff.
Pass the reviewer:
git -C $target_repo_dir diff --cached.triage_result (so the reviewer knows what the root cause is
supposed to be).allow_skip (always false when this step runs).Instruct the reviewer to reject the diff (return REQUEST_CHANGES) if
any of the following appears anywhere in added lines:
@skipIfXpu / @skipXPU / @skipCUDAIf / @skipMPS /
@unittest.skip / @unittest.skipIf / @pytest.mark.skip /
@pytest.mark.skipif / @pytest.mark.xfail /
@expectedFailureXPU / @expectedFailureCUDA / @expectedFailureMPS
/ @expectedFailure decorator on a test.DecorateInfo(unittest.skip, ...) / DecorateInfo(skipIfXpu, ...)
/ DecorateInfo(unittest.expectedFailure, ...) entry in an
OpInfo / ModuleInfo skips/decorators list.xfail(...) / skip(...) entry in an
instantiate_device_type_tests skip dict for XPU.raise unittest.SkipTest(...) / self.skipTest(...) inserted
into a previously-running test to short-circuit it on XPU.atol / rtol on assertEqual (or any tolerance-carrying
assertion) by more than an order of magnitude, when the diff has no
quantitative justification for the new value in a comment.set_rng_seed(...) / torch.manual_seed(...) /
random.seed(...) inserted into a previously-random test purely to
dodge a failure region.try / except Exception: pass (or equivalent) wrapping the
call that used to fail.Existing skips being removed by the diff are fine — that is a legitimate root-cause fix in the "stale test expectation" category. The rule only fires on added skip-shaped constructs.
The reviewer returns one of:
APPROVE — no skip-shaped workaround found. Continue to Step 4
(output).REQUEST_CHANGES — cite each offending hunk (file + line + which
rule it matched). The implementer MUST address every citation
(either replace the workaround with a real root-cause fix or, if no
root-cause fix is possible within this run's scope, unstage the
offending change and return NEEDS_HUMAN(reason=skip_guard_rejected)
to the orchestrator with the reviewer's citations attached).Do not loop this step more than once. If a second run of Step 3.5
still returns REQUEST_CHANGES, unstage the offending change and
return NEEDS_HUMAN(reason=skip_guard_rejected) — that is a signal the
fix cannot be produced without a workaround and a human should take it.
This step is intentionally narrower than the orchestrator's Stage 5.5
review. Stage 5.5 checks the entire diff for correctness, minimalism,
and root-cause alignment; Step 3.5 checks only for the specific class of
"hide the failure instead of fixing it" workarounds that
allow_skip=false is meant to forbid. Both run; they do not replace
each other.
Return to the orchestrator a report (markdown block plus JSON block).
The skill does not commit, push, or open PRs — the caller consumes
stdout and decides what to do, per the pattern established by
issue-triage, fix-reproduce, and fix-root-cause.
Include the <!-- agent:implement --> marker on the first line of the
markdown block so a downstream caller can locate its own previous
implement comment (if any) and update it in place. Comment location and
update is the caller's responsibility.
<!-- agent:implement -->
## Implement Result
- **Target repo:** <pytorch | torch-xpu-ops>
- **Analyzed at:** <target_repo>@<short_sha>
- **What I changed:** <per file: what changed in it, with the line count>
- **Why:** <one sentence connecting each change to the triage root cause; must cite file:line for every upstream/CUDA comparison — see the hard rule below>
- **Skip added:** <yes (tracking: intel/torch-xpu-ops#N, url: <url>) | no>
- **Ready for verify:** <yes | no>
```diff
<git diff --cached, or the key hunks when the full diff would blow the
65000-char comment limit>
```
*Automated by fix-implement.*{
"target_repo": "pytorch or torch-xpu-ops",
"analyzed_sha": "<full 40-char sha inherited from triage_result>",
"changed_files": ["path/to/file1.py", "src/ATen/native/xpu/Foo.cpp"],
"skip_added": false,
"tracking_issue": null,
"allow_skip": false,
"covers": [],
"ready_for_verify": true,
"verdict": "READY or NEEDS_HUMAN",
"reason": "<enumerated reason code, see below>",
"reason_detail": "one-line human-readable detail"
}target_repo / analyzed_sha — echo from triage_result so
downstream stages have a self-contained record without re-reading
triage output.changed_files — list of paths (relative to target_repo_dir) that
are staged. fix-verify reads this to decide whether a C++/SYCL
rebuild is required. Must equal git -C $target_repo_dir diff --cached --name-only at output time. Re-run that command immediately
before emitting the JSON block; do not cache a pre-Step-3.5 file list,
because Step 3.5's reviewer may have unstaged offending files.covers — other batch entries this patch also fixes, confirmed by
re-running them against the staged fix (see issue-handler, "One fix
may cover several sub-items"). Node ids or issue numbers; empty on the
single-bug path. They are not fixed again.skip_added — true only when this run added a new skip decorator
under allow_skip=true. Removing a stale skip is NOT skip_added.tracking_issue — issue URL from the "Add a new skip" recipe in
Step 2 when skip_added=true; null otherwise. The orchestrator
reads this for the "Skipped (with tracking issue)" rows of its
fan-out report.allow_skip — echo the input flag verbatim, so a reviewer can tell
by looking at the output alone whether Step 3.5 ran.ready_for_verify — true when Step 3.5 (if it ran) returned
APPROVE and staged changes exist; false if the implementer
decided to bail out with NEEDS_HUMAN (in which case the
orchestrator should not call fix-verify).verdict — READY (staged diff exists, fix-verify may run) or
NEEDS_HUMAN (see reason).reason valuesOn verdict=READY: ok.
On verdict=NEEDS_HUMAN:
skip_outside_target_repo — the skip decorator that would need to be
added lives outside target_repo_dir; see Step 2's "Skip operations".skip_guard_rejected — Step 3.5's subagent reviewer returned
REQUEST_CHANGES twice, or the implementer chose to bail out rather
than fight the reviewer.no_fix_possible — the fix strategy in triage_result is not
implementable without either widening the diff outside target_repo
or adding a skip when allow_skip=false.other — fallback; put full explanation in reason_detail.Contract: this leaf leaves the changes staged (git add) and
does not commit — leaves never commit, branch, or push. What the caller
does with the staged diff (commit it onto a branch, or hand the staged
diff straight to its workflow) is the caller's business; this skill only
stages.
The "Why" line written here will be read by reviewers as factual statements. Before writing any of the following phrases, you MUST look up the specific file and line number that backs it up:
If you cannot find a specific file:line to cite, do not write the
phrase. Replace it with a direct statement of the observable fact,
e.g.:
| Instead of... | Write... |
|---|---|
| "consistent with upstream's bcomplex32 handling" | "BComplex32 comparison (isclose/mul) raises NotImplementedError in the nightly wheel (pytorch warns 'BComplex32 support is experimental')" |
| "upstream already skips this" | "upstream test_ops.py:595 skips {bfloat16, bcomplex32} inputs in _ref_test_helper" |
A "Why" that contains an unsubstantiated upstream comparison is worse than one that only cites what you directly observed — it inflates reviewer confidence in a claim that was never verified.
allow_skip=false.target_repo_dir — including when the skip
or fix "belongs" in the other repo. That diff cannot be committed or
read back by the orchestrator; return NEEDS_HUMAN instead.allow_skip=false, Step 3.5 (skip-guard reviewer subagent) is
MANDATORY before returning to the orchestrator. Do not skip it, do
not run it inline in your own context.third_party/xpu.txt. It is a submodule pin managed by
build tooling; staging it is never part of a bug fix, regardless of
domain (xpu-kernel / inductor / upstream-pytorch).git rebase origin/main) instead.git add); the workflow that invoked it reads the staged diff
after fix-verify passes and exports it as a patch (no branch, no
push) for a human to apply.© intel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/fix-implement of intel/torch-xpu-ops.
Open the folder on GitHubat commit 0187b3b
Fix Implement next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Fix Implement this skillintel/torch-xpu-ops | 115 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| The Art of Debuggingstas00/the-art-of-debugging | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | |
| Ascendcascend-ai-coding/awesome-ascend-skills | 174 | — | ~3.5k | Automated safety check: Pass | None | |
| Vllm Pytorch CI Triagepytorch/test-infra | 113 | — | ~2.3k | Automated safety check: Pass | Custom licence | |
| Vllm Upstream Deduppytorch/test-infra | 113 | — | ~1.3k | Automated safety check: Pass | Custom licence | |
| ExecuTorch Binary Size Reductionpytorch/executorch | 5.1k | — | ~793 | Automated safety check: Pass | Custom licence |
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
ascend-ai-coding/awesome-ascend-skills
End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project.
pytorch/test-infra
Root-cause a vLLM torch-nightly CI regression report. An agent skill from pytorch/test-infra.
pytorch/test-infra
Review vLLM-routed torch-nightly root causes against existing upstream vLLM issues using read-only search results, and emit a validated upstream-checks artifact for the filer.
pytorch/executorch
Measures and shrinks the ExecuTorch runtime binary by building a size test, analyzing it with bloaty and landing each reduction as its own pull request.
pytorch/executorch
Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.
intel/torch-xpu-ops
Select the Intel GPU device to use when a system has multiple Intel GPU devices.
intel/torch-xpu-ops
Check PyTorch ciflow/xpu (xpu.yml) on the main branch, collect the failing XPU test cases from the most recent completed run(s), analyze the ROOT CAUSE of each failure with AI, and produce a list…
intel/torch-xpu-ops
Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.
intel/torch-xpu-ops
Review pull requests for XPU operator or backend code. An agent skill from intel/torch-xpu-ops.
intel/torch-xpu-ops
Guide users through creating Agent Skills for Claude Code. An agent skill from intel/torch-xpu-ops.
intel/torch-xpu-ops
Read the evidence a nightly UT run produced, decide which failures share a root cause and which are machine breakage rather than product bugs, and write one issue draft per root cause to drafts.json.
Works with
Categories
A skill your agent uses when asked to implement a fix for a root-caused failure, apply a proposed patch, or produce a staged code change from triage output. Fix Implement is an agent skill from intel/torch-xpu-ops, published by the product's own GitHub organization. Use when asked to implement a fix for a root-caused failure, apply a proposed patch, or produce a staged code change from triage output.
Fix Implement fits situations like: asked to implement a fix for a root-caused failure; apply a proposed patch; produce a staged code change from triage output.
Run `npx skills add intel/torch-xpu-ops --skill fix-implement -a claude-code`. Or copy the skill folder (.claude/skills/fix-implement in intel/torch-xpu-ops) into .claude/skills/fix-implement in your project. Claude Code loads it when a task matches its description.
Run `npx skills add intel/torch-xpu-ops --skill fix-implement -a codex`. Or copy the skill folder (.claude/skills/fix-implement in intel/torch-xpu-ops) into .agents/skills/fix-implement in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intel/torch-xpu-ops --skill fix-implement -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fix-implement, .gemini/skills/fix-implement, .github/skills/fix-implement and .opencode/skills/fix-implement in your project.
Going by SKILL.md and its folder, Fix Implement needs the command-line tools its instructions call (git and gh).
SKILL.md contains no URLs. Its commands use git and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Fix Implement is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Fix Implement: The Art of Debugging (stas00/the-art-of-debugging, 1.7k stars), Ascendc (ascend-ai-coding/awesome-ascend-skills, 174 stars), Vllm Pytorch CI Triage (pytorch/test-infra, 113 stars) and Vllm Upstream Dedup (pytorch/test-infra, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
intel (a GitHub organization, an official publisher) maintains it in intel/torch-xpu-ops, which has 115 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 6, 2026.
Source: intel/torch-xpu-ops on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.