Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently…
Install the "coreweave-gpu-node-forensics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-node-forensics into .claude/skills/coreweave-gpu-node-forensics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-node-forensics", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "coreweave-gpu-node-forensics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-node-forensics into .agents/skills/coreweave-gpu-node-forensics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-node-forensics", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "coreweave-gpu-node-forensics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-node-forensics into .cursor/skills/coreweave-gpu-node-forensics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-node-forensics", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "coreweave-gpu-node-forensics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-node-forensics into .gemini/skills/coreweave-gpu-node-forensics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-node-forensics", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "coreweave-gpu-node-forensics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-node-forensics into .github/skills/coreweave-gpu-node-forensics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-node-forensics", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "coreweave-gpu-node-forensics" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-node-forensics into .opencode/skills/coreweave-gpu-node-forensics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-node-forensics", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
coreweave-gpu-node-forensics
GitHub stars
2.8k
Token cost
~3.2k tokens
SKILL.md length
1,418 words
Files
8 (incl. scripts, references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT
At a glance
Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently…
Works in 6 steps: Capture the evidence → Run the deterministic triage → Read the verdict and act on the action → …
A GPU throws an Xid error
SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
Runs Python scripts from its folder; calls python3, kubectl and jq
What it does
Coreweave GPU Node Forensics is an agent skill from jeremylongshore/tons-of-skills-marketplace. Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently kill a multi-day training run. Use when a GPU throws an Xid error, a node "fell off the bus", a training run stalls or NCCL hangs on one rank, or you need to know whether to replace, reset, or just reschedule. Trigger with "xid error", "gpu fell off the bus", "coreweave gpu dead", "should I RMA this GPU", "gpu node…
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `ARD.md`, `PRD.md` and `eval-spec.yaml`). Compatibility notes: Designed for Claude Code
It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with NVIDIA AI Platform. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.
When your agent uses it
A GPU throws an Xid error
A node fell off the bus
A training run stalls
NCCL hangs on one rank
Example prompts
“fell off the bus”
“xid error”
“gpu fell off the bus”
“/coreweave-gpu-node-forensics”
Requirements
Python 3
Compatibility (from SKILL.md): Designed for Claude Code
Read from SKILL.md and the folder at commit 23ea8d4. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves these tools, so the agent can use them without asking each time:
Read
Bash(python3:*)
Bash(nvidia-smi -q:*)
Bash(kubectl get:*)
Bash(dmesg:*)
From allowed-tools in the SKILL.md frontmatter.
Runs code
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3
kubectl
jq
From the folder's file list and the shell code blocks in SKILL.md.
Network
Links to these hosts (documentation or services it may open):
docs.nvidia.com
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Compatibility
Designed for Claude Code
From compatibility in the SKILL.md frontmatter.
Context cost
Coreweave GPU Node Forensics loads about 3.2k tokens when it runs, and up to ~6.2k if it reads all its reference files. Until then it costs about 138 tokens; SKILL.md has 1,418 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~138
When it runs· the whole SKILL.md, loaded when a task matches
~3.2k
With references· SKILL.md plus every file in references/, read only if the agent opens them
~6.2k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Download SKILL.mdSave it as .claude/skills/coreweave-gpu-node-forensics/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
coreweave-gpu-node-forensics
description
Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule
vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg /
nvidia-smi blob, so a bad card does not silently kill a multi-day training run.
Use when a GPU throws an Xid error, a node "fell off the bus", a training run
stalls or NCCL hangs on one rank, or you need to know whether to replace,
reset, or just reschedule. Trigger with "xid error", "gpu fell off the bus",
"coreweave gpu dead", "should I RMA this GPU", "gpu node triage".
Community-contributed. Not affiliated with, endorsed by, or sponsored by
CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
"NVIDIA" and "Xid" are trademarks of NVIDIA Corporation; Xid semantics are
cited from NVIDIA's public documentation.
Triages a dead or degraded GPU on a CoreWeave node in seconds and returns one
grounded move — reschedule, reset-gpu, reboot-node, rma, watch, or
app-bug-not-hardware — from an NVIDIA Xid code or a pasted dmesg /
nvidia-smi blob. The classification is deterministic: a bundled script does the
mapping so the agent never guesses whether a card is dead, degraded, or fine.
Overview
A single bad GPU can kill a 64-GPU, multi-day training run — thousands of dollars
and days of wall-clock gone — because one rank stalls the whole collective. The
expensive mistakes are triage mistakes: RMAing a healthy card for an app bug,
restarting a job onto a GPU whose memory error was uncontained, or manually
uncordoning a node the lifecycle controller is trying to replace. This skill kills
that ambiguity.
The decision logic is grounded in the NVIDIA Xid error catalog
(https://docs.nvidia.com/deploy/xid-errors/) and CoreWeave's node-lifecycle /
cordon behavior. The math-of-the-matter — which Xid means what, and how
row-remapper state overrides it — lives in scripts/triage.py as a table the LLM
does not get to re-litigate. Deep domain knowledge (the full code→action table,
the row-remap decision, the cordon rules) loads from references/ on demand.
The headline is the Xid 94-vs-95 split. A contained memory error (94) cost
you one job restart on a healthy node; an uncontained one (95) means the GPU
could not isolate the fault and everything it touched is suspect. Getting that one
bit wrong is the difference between a 30-second reschedule and a run that quietly
trained on corrupt gradients. The script decides it; the skill never eyeballs it.
This skill is diagnostic, not destructive: it recommends the cordon / drain /
reset / RMA next-step but its tools are scoped read-only (nvidia-smi -q,
kubectl get, dmesg) — it never runs a reset, a reboot, or an uncordon itself.
Prerequisites
The failure evidence. Either an Xid number, or a pasted dmesg /
nvidia-smi dump. The skill works from a paste alone — no live cluster access
required — which is the common case (an operator pastes what the run's logs
showed).
Optional live access for corroboration: kubectl context on the CoreWeave
cluster (read-only is enough), and nvidia-smi on the node. If neither is
available the skill still triages from the paste.
python3 to run the deterministic classifier (scripts/triage.py, stdlib
only — no dependencies).
No secrets are handled. All commands are read-only queries.
Instructions
The pipeline is capture → classify → act. The classifier is authoritative for
the verdict; references/ supplies the "why" when a case needs depth.
Step 1: Capture the evidence
If the user has not already pasted it, ask for (or read) the fault signal. The two
richest sources:
Feed the Xid code, or the whole blob, to the classifier. Do not classify by
hand — the script owns the Xid→action mapping and the row-remap override.
bash
# From an Xid code:
python3 "${CLAUDE_SKILL_DIR}/scripts/triage.py" --xid 95
# With row-remapper state (Xid 63/64 or a DBE):
python3 "${CLAUDE_SKILL_DIR}/scripts/triage.py" --xid 63 --pending yes
python3 "${CLAUDE_SKILL_DIR}/scripts/triage.py" --xid 48 --remap-failure yes
# From a pasted dmesg / nvidia-smi blob (file or stdin):
python3 "${CLAUDE_SKILL_DIR}/scripts/triage.py" --blob /path/to/dmesg.txt
dmesg -T | grep -i xid | python3 "${CLAUDE_SKILL_DIR}/scripts/triage.py"
The script emits a VERDICT block plus a JSON object
({classification, severity, action, why, next_command, cordon_rule?, unverified?}).
Add --json for machine-readable output only.
When a blob carries several Xids, the most severe one governs the move and the
rest are reported as co_occurring_xids — a co-occurring app-side Xid 43 next to a
hardware Xid 79 does not soften the "reboot the node" verdict.
Step 3: Read the verdict and act on the action
Present the action and the next_command to the user in plain language. The six
actions and what each means live in
${CLAUDE_SKILL_DIR}/references/xid-triage-table.md.
Load it when the user wants the full table or asks about an Xid not in the summary.
reschedule — restart the failed rank; the node stays in service.
reset-gpu — drain the GPU and reset it; re-run any job that shared it.
reboot-node — the card is off the bus; only a bare-metal reboot returns it.
rma — terminal hardware fault; the card must be replaced.
watch — correctable / trending; monitor, do not act yet.
app-bug-not-hardware — the app faulted, not the GPU. Do NOT RMA.
Step 4: The 94-vs-95 headline (load-bearing — get this right)
If the Xid is 94 or 95, state the containment explicitly, because the two look
almost identical in the logs and lead to opposite actions:
Xid 94 (CONTAINED) → reschedule. The error was isolated to the app's
context; the GPU and node are healthy. Restart the job. Do not reset or RMA.
Xid 95 (UNCONTAINED) → reset-gpu. The GPU could not isolate it; every
context it touched is suspect. Drain, reset, and re-run anything that shared the
card — otherwise the run may continue on corrupt state.
For any ECC/DBE or row-remap Xid, the nvidia-smi -q -d ROW_REMAPPER fields
override the base verdict. Pass them in (--pending, --remap-failure) and let
the script decide. The rule:
Whenever the action is a hardware move (reset-gpu, reboot-node, rma), the
verdict carries a cordon_rule. Surface it verbatim:
Never manually uncordon a CoreWeave health cordon — the node-lifecycle
controller owns cordon/uncordon and is driving remediation. Uncordoning
re-admits the run onto known-bad hardware and races the controller.
A verdict — the deterministic {classification, severity, action, why, next_command} for the governing signal, presented in plain language with the
single next command an operator runs.
The 94-vs-95 call stated explicitly when either fires — contained →
reschedule, uncontained → reset — never conflated.
The cordon rule attached to every hardware action, verbatim, so nobody
uncordons a node the controller is replacing.
Co-occurring Xids listed when a blob carried several, with the most-severe
one named as governing.
[unverified] hedges surfaced honestly — e.g. CoreWeave's exact auto-RMA
thresholds and health-cordon taint strings are not publicly published, and the
skill says so rather than inventing a number.
Error Handling
Situation
Cause
Response
Unmapped Xid
Xid not in the triage table
Script returns a conservative reset-gpu default, flags it [unverified], and points at the NVIDIA catalog. Do not RMA on an unmapped Xid.
Xid 94 vs 95 ambiguity
Both are "ECC error" in casual logs
Never infer from prose — use the exact code. 94 = contained (reschedule), 95 = uncontained (reset).
Row-remap without an Xid
Only nvidia-smi pasted
Script reads Remapping Failure Occurred / Pending and decides from those alone.
App-side Xid mistaken for hardware
Xid 13/31/43/45
Classed app-bug-not-hardware; the fix is compute-sanitizer on the job, never an RMA.
Multiple Xids in one blob
Cascade (e.g. 43 then 79)
Most-severe governs; others listed as co_occurring_xids.
Empty / no signal
Blob has no Xid, remap, or thermal line
Script returns watch and asks for fresh dmesg + nvidia-smi -q evidence.
Thermal throttle, not a fault
HW/SW Thermal Slowdown : Active
Classed watch — healthy but hot; investigate cooling, not the card.
Tempted to uncordon a stalled node
Health cordon owned by the controller
Refuse. Surface the cordon rule; let remediation run.
Examples
Example 1: "Xid 94 on one rank — do I need to replace the GPU?"
text
VERDICT: Xid 94 — Contained ECC/memory error [severity: HIGH]
ACTION: reschedule
WHY: The error was CONTAINED to the faulting application's context — the GPU
and node are healthy. Just reschedule/restart the job; no reset or RMA needed.
NEXT: Restart or reschedule the failed rank; the node stays in service.
No replacement. The containment did its job — restart the rank and keep the node.
Example 2: "Xid 95 — same run, different node"
text
VERDICT: Xid 95 — Uncontained ECC/memory error [severity: CRITICAL]
ACTION: reset-gpu
WHY: The error was UNCONTAINED — the GPU could not isolate it, so every context
it touched is suspect. Drain the GPU and reset it; treat all in-flight work as corrupt.
NEXT: Cordon + drain, GPU-reset (nvidia-smi -r) or node reset; re-run any job that shared this GPU.
CORDON: Do NOT manually uncordon a CoreWeave health cordon — the node-lifecycle
controller owns cordon/uncordon. Let it drain and replace the node.
Opposite of Example 1 despite looking identical in the logs: drain, reset, and
re-run the shared work — do not just restart.
Example 3: "GPU fell off the bus"
python3 scripts/triage.py --xid 79 → reboot-node. The card is unreachable on
the PCIe bus and returns only after a bare-metal reboot; expect the controller to
cordon and reboot — let it, and RMA only if it recurs after the reboot.
Example 4: "Xid 63 with a remap failure"
python3 scripts/triage.py --xid 63 --remap-failure yes → rma. The remapper
tried to swap in a spare row and physically could not — terminal. Cordon, drain,
open the RMA. (The script flags that CoreWeave's exact auto-RMA threshold is
[unverified].)
Example 5: "Xid 43 killed my job"
python3 scripts/triage.py --xid 43 → app-bug-not-hardware. Software-induced
fault — the fix is in the application (run it under compute-sanitizer), not an RMA.
Pulling the card would waste healthy hardware and never fix the job.
Coreweave GPU Node Forensics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Coreweave GPU Node Forensics compared with similar skills
Skill
Stars
Used in
Tokens
Auto-check
Licence
Repo updated
Coreweave GPU Node Forensics this skilljeremylongshore/tons-of-skills-marketplace
Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…
Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently…. Coreweave GPU Node Forensics is an agent skill from jeremylongshore/tons-of-skills-marketplace. Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently kill a multi-day training run.
When should I use Coreweave GPU Node Forensics?
Coreweave GPU Node Forensics fits situations like: A GPU throws an Xid error; A node fell off the bus; A training run stalls; NCCL hangs on one rank.
How do I install Coreweave GPU Node Forensics in Claude Code?
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a claude-code`. Or copy the skill folder (skills/.curated/coreweave-gpu-node-forensics in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/coreweave-gpu-node-forensics in your project. Claude Code loads it when a task matches its description.
How do I install Coreweave GPU Node Forensics in Codex?
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a codex`. Or copy the skill folder (skills/.curated/coreweave-gpu-node-forensics in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/coreweave-gpu-node-forensics in your project. Codex loads it when a task matches its description.
Can I use Coreweave GPU Node Forensics in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coreweave-gpu-node-forensics, .gemini/skills/coreweave-gpu-node-forensics, .github/skills/coreweave-gpu-node-forensics and .opencode/skills/coreweave-gpu-node-forensics in your project.
What does Coreweave GPU Node Forensics need to run?
Going by SKILL.md and its folder, Coreweave GPU Node Forensics needs Python for the scripts in its folder and the command-line tools its instructions call (python3, kubectl and jq). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Bash(python3:*), Bash(nvidia-smi -q:*), Bash(kubectl get:*), Bash(dmesg:*). Compatibility (from SKILL.md): Designed for Claude Code.
Does Coreweave GPU Node Forensics access the network?
SKILL.md names 1 domain. As links in the text: docs.nvidia.com. This is read from the text; nothing was executed.
Is Coreweave GPU Node Forensics safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
What licence does Coreweave GPU Node Forensics use?
Coreweave GPU Node Forensics is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Coreweave GPU Node Forensics use?
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
What are the alternatives to Coreweave GPU Node Forensics?
Skills that share tags, products or a category with Coreweave GPU Node Forensics: Fla Triton To Gluon (fla-org/flash-linear-attention, 5.8k stars), DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars), DGX Spark Training Gotchas (wshobson/agents, 40k stars) and Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Coreweave GPU Node Forensics?
jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,821 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 8, 2026.