Agent skill

Coreweave GPU Node Forensics

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently…

MITAuto-check passedAI & LLM Engineering

Install Coreweave GPU Node Forensics

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-gpu-node-forensics --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/coreweave-gpu-node-forensics .claude/skills/coreweave-gpu-node-forensics && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
coreweave-gpu-node-forensics
GitHub stars
2.8k
Token cost
~3.2k tokens
SKILL.md length
1,418 words
Files
8 (incl. scripts, references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently…

  • Works in 6 steps: Capture the evidence → Run the deterministic triage → Read the verdict and act on the action → …
  • A GPU throws an Xid error
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
  • Runs Python scripts from its folder; calls python3, kubectl and jq

What it does

Coreweave GPU Node Forensics is an agent skill from jeremylongshore/tons-of-skills-marketplace. Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently kill a multi-day training run. Use when a GPU throws an Xid error, a node "fell off the bus", a training run stalls or NCCL hangs on one rank, or you need to know whether to replace, reset, or just reschedule. Trigger with "xid error", "gpu fell off the bus", "coreweave gpu dead", "should I RMA this GPU", "gpu node…

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `ARD.md`, `PRD.md` and `eval-spec.yaml`). Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with NVIDIA AI Platform. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • A GPU throws an Xid error
  • A node fell off the bus
  • A training run stalls
  • NCCL hangs on one rank

Example prompts

  • “fell off the bus”
  • “xid error”
  • “gpu fell off the bus”
  • “/coreweave-gpu-node-forensics”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Bash(python3:*), Bash(nvidia-smi -q:*), Bash(kubectl get:*), Bash(dmesg:*)

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Capture the evidence
  2. Run the deterministic triage
  3. Read the verdict and act on the action
  4. The 94-vs-95 headline (load-bearing — get this right)
  5. Row-remap (Xid 63 / 64 / 48) — routine vs terminal
  6. The cordon hard rule (never violate)

What it can do on your machine

Read from SKILL.md and the folder at commit 23ea8d4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash(python3:*)
    • Bash(nvidia-smi -q:*)
    • Bash(kubectl get:*)
    • Bash(dmesg:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • kubectl
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Coreweave GPU Node Forensics loads about 3.2k tokens when it runs, and up to ~6.2k if it reads all its reference files. Until then it costs about 138 tokens; SKILL.md has 1,418 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~138
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit 23ea8d4, republished under its MIT licence (© jeremylongshore). 1,418 words, ~3,207 tokens.

Download SKILL.mdSave it as .claude/skills/coreweave-gpu-node-forensics/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
coreweave-gpu-node-forensics
description
Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently kill a multi-day training run. Use when a GPU throws an Xid error, a node "fell off the bus", a training run stalls or NCCL hangs on one rank, or you need to know whether to replace, reset, or just reschedule. Trigger with "xid error", "gpu fell off the bus", "coreweave gpu dead", "should I RMA this GPU", "gpu node triage".
allowed-tools
Read, Bash(python3:*), Bash(nvidia-smi -q:*), Bash(kubectl get:*), Bash(dmesg:*)
compatibility
Designed for Claude Code
version
1.11.0
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT
tags
saas, coreweave, gpu-cloud, reliability, xid

CoreWeave GPU Node Forensics

Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc. "NVIDIA" and "Xid" are trademarks of NVIDIA Corporation; Xid semantics are cited from NVIDIA's public documentation.

Triages a dead or degraded GPU on a CoreWeave node in seconds and returns one grounded move — reschedule, reset-gpu, reboot-node, rma, watch, or app-bug-not-hardware — from an NVIDIA Xid code or a pasted dmesg / nvidia-smi blob. The classification is deterministic: a bundled script does the mapping so the agent never guesses whether a card is dead, degraded, or fine.

Overview

A single bad GPU can kill a 64-GPU, multi-day training run — thousands of dollars and days of wall-clock gone — because one rank stalls the whole collective. The expensive mistakes are triage mistakes: RMAing a healthy card for an app bug, restarting a job onto a GPU whose memory error was uncontained, or manually uncordoning a node the lifecycle controller is trying to replace. This skill kills that ambiguity.

The decision logic is grounded in the NVIDIA Xid error catalog (https://docs.nvidia.com/deploy/xid-errors/) and CoreWeave's node-lifecycle / cordon behavior. The math-of-the-matter — which Xid means what, and how row-remapper state overrides it — lives in scripts/triage.py as a table the LLM does not get to re-litigate. Deep domain knowledge (the full code→action table, the row-remap decision, the cordon rules) loads from references/ on demand.

The headline is the Xid 94-vs-95 split. A contained memory error (94) cost you one job restart on a healthy node; an uncontained one (95) means the GPU could not isolate the fault and everything it touched is suspect. Getting that one bit wrong is the difference between a 30-second reschedule and a run that quietly trained on corrupt gradients. The script decides it; the skill never eyeballs it.

This skill is diagnostic, not destructive: it recommends the cordon / drain / reset / RMA next-step but its tools are scoped read-only (nvidia-smi -q, kubectl get, dmesg) — it never runs a reset, a reboot, or an uncordon itself.

Prerequisites

  • The failure evidence. Either an Xid number, or a pasted dmesg / nvidia-smi dump. The skill works from a paste alone — no live cluster access required — which is the common case (an operator pastes what the run's logs showed).
  • Optional live access for corroboration: kubectl context on the CoreWeave cluster (read-only is enough), and nvidia-smi on the node. If neither is available the skill still triages from the paste.
  • python3 to run the deterministic classifier (scripts/triage.py, stdlib only — no dependencies).

No secrets are handled. All commands are read-only queries.

Instructions

The pipeline is capture → classify → act. The classifier is authoritative for the verdict; references/ supplies the "why" when a case needs depth.

Step 1: Capture the evidence

If the user has not already pasted it, ask for (or read) the fault signal. The two richest sources:

bash
dmesg -T | grep -i xid                       # the Xid line(s) with timestamps
nvidia-smi -q -d ROW_REMAPPER,ECC,PERFORMANCE # remap state, ECC counts, throttle reasons

On a CoreWeave node you can also check who owns any cordon before acting:

bash
kubectl get node NODE -o json | jq '{unschedulable: .spec.unschedulable, taints: .spec.taints}'
Step 2: Run the deterministic triage

Feed the Xid code, or the whole blob, to the classifier. Do not classify by hand — the script owns the Xid→action mapping and the row-remap override.

bash
# From an Xid code:
python3 "${CLAUDE_SKILL_DIR}/scripts/triage.py" --xid 95

# With row-remapper state (Xid 63/64 or a DBE):
python3 "${CLAUDE_SKILL_DIR}/scripts/triage.py" --xid 63 --pending yes
python3 "${CLAUDE_SKILL_DIR}/scripts/triage.py" --xid 48 --remap-failure yes

# From a pasted dmesg / nvidia-smi blob (file or stdin):
python3 "${CLAUDE_SKILL_DIR}/scripts/triage.py" --blob /path/to/dmesg.txt
dmesg -T | grep -i xid | python3 "${CLAUDE_SKILL_DIR}/scripts/triage.py"

The script emits a VERDICT block plus a JSON object ({classification, severity, action, why, next_command, cordon_rule?, unverified?}). Add --json for machine-readable output only.

When a blob carries several Xids, the most severe one governs the move and the rest are reported as co_occurring_xids — a co-occurring app-side Xid 43 next to a hardware Xid 79 does not soften the "reboot the node" verdict.

Step 3: Read the verdict and act on the action

Present the action and the next_command to the user in plain language. The six actions and what each means live in ${CLAUDE_SKILL_DIR}/references/xid-triage-table.md. Load it when the user wants the full table or asks about an Xid not in the summary.

  • reschedule — restart the failed rank; the node stays in service.
  • reset-gpu — drain the GPU and reset it; re-run any job that shared it.
  • reboot-node — the card is off the bus; only a bare-metal reboot returns it.
  • rma — terminal hardware fault; the card must be replaced.
  • watch — correctable / trending; monitor, do not act yet.
  • app-bug-not-hardware — the app faulted, not the GPU. Do NOT RMA.
Step 4: The 94-vs-95 headline (load-bearing — get this right)

If the Xid is 94 or 95, state the containment explicitly, because the two look almost identical in the logs and lead to opposite actions:

  • Xid 94 (CONTAINED) → reschedule. The error was isolated to the app's context; the GPU and node are healthy. Restart the job. Do not reset or RMA.
  • Xid 95 (UNCONTAINED) → reset-gpu. The GPU could not isolate it; every context it touched is suspect. Drain, reset, and re-run anything that shared the card — otherwise the run may continue on corrupt state.
Step 5: Row-remap (Xid 63 / 64 / 48) — routine vs terminal

For any ECC/DBE or row-remap Xid, the nvidia-smi -q -d ROW_REMAPPER fields override the base verdict. Pass them in (--pending, --remap-failure) and let the script decide. The rule:

  • Remapping Failure Occurred: Yes → rma (terminal — sparing failed).
  • Pending: Yes (no failure) → reset-gpu (routine — applies on reset).

Full logic + how to read the histogram: ${CLAUDE_SKILL_DIR}/references/row-remap-decision.md.

Show full SKILL.md (574 more words)Show less
Step 6: The cordon hard rule (never violate)

Whenever the action is a hardware move (reset-gpu, reboot-node, rma), the verdict carries a cordon_rule. Surface it verbatim:

Never manually uncordon a CoreWeave health cordon — the node-lifecycle controller owns cordon/uncordon and is driving remediation. Uncordoning re-admits the run onto known-bad hardware and races the controller.

Determine cordon provenance before touching schedulability. Full guidance: ${CLAUDE_SKILL_DIR}/references/cordon-rules.md.

Output

  • A verdict — the deterministic {classification, severity, action, why, next_command} for the governing signal, presented in plain language with the single next command an operator runs.
  • The 94-vs-95 call stated explicitly when either fires — contained → reschedule, uncontained → reset — never conflated.
  • The cordon rule attached to every hardware action, verbatim, so nobody uncordons a node the controller is replacing.
  • Co-occurring Xids listed when a blob carried several, with the most-severe one named as governing.
  • [unverified] hedges surfaced honestly — e.g. CoreWeave's exact auto-RMA thresholds and health-cordon taint strings are not publicly published, and the skill says so rather than inventing a number.

Error Handling

SituationCauseResponse
Unmapped XidXid not in the triage tableScript returns a conservative reset-gpu default, flags it [unverified], and points at the NVIDIA catalog. Do not RMA on an unmapped Xid.
Xid 94 vs 95 ambiguityBoth are "ECC error" in casual logsNever infer from prose — use the exact code. 94 = contained (reschedule), 95 = uncontained (reset).
Row-remap without an XidOnly nvidia-smi pastedScript reads Remapping Failure Occurred / Pending and decides from those alone.
App-side Xid mistaken for hardwareXid 13/31/43/45Classed app-bug-not-hardware; the fix is compute-sanitizer on the job, never an RMA.
Multiple Xids in one blobCascade (e.g. 43 then 79)Most-severe governs; others listed as co_occurring_xids.
Empty / no signalBlob has no Xid, remap, or thermal lineScript returns watch and asks for fresh dmesg + nvidia-smi -q evidence.
Thermal throttle, not a faultHW/SW Thermal Slowdown : ActiveClassed watch — healthy but hot; investigate cooling, not the card.
Tempted to uncordon a stalled nodeHealth cordon owned by the controllerRefuse. Surface the cordon rule; let remediation run.

Examples

Example 1: "Xid 94 on one rank — do I need to replace the GPU?"
text
VERDICT: Xid 94 — Contained ECC/memory error [severity: HIGH]
ACTION:  reschedule
WHY:     The error was CONTAINED to the faulting application's context — the GPU
         and node are healthy. Just reschedule/restart the job; no reset or RMA needed.
NEXT:    Restart or reschedule the failed rank; the node stays in service.

No replacement. The containment did its job — restart the rank and keep the node.

Example 2: "Xid 95 — same run, different node"
text
VERDICT: Xid 95 — Uncontained ECC/memory error [severity: CRITICAL]
ACTION:  reset-gpu
WHY:     The error was UNCONTAINED — the GPU could not isolate it, so every context
         it touched is suspect. Drain the GPU and reset it; treat all in-flight work as corrupt.
NEXT:    Cordon + drain, GPU-reset (nvidia-smi -r) or node reset; re-run any job that shared this GPU.
CORDON:  Do NOT manually uncordon a CoreWeave health cordon — the node-lifecycle
         controller owns cordon/uncordon. Let it drain and replace the node.

Opposite of Example 1 despite looking identical in the logs: drain, reset, and re-run the shared work — do not just restart.

Example 3: "GPU fell off the bus"

python3 scripts/triage.py --xid 79 → reboot-node. The card is unreachable on the PCIe bus and returns only after a bare-metal reboot; expect the controller to cordon and reboot — let it, and RMA only if it recurs after the reboot.

Example 4: "Xid 63 with a remap failure"

python3 scripts/triage.py --xid 63 --remap-failure yes → rma. The remapper tried to swap in a spare row and physically could not — terminal. Cordon, drain, open the RMA. (The script flags that CoreWeave's exact auto-RMA threshold is [unverified].)

Example 5: "Xid 43 killed my job"

python3 scripts/triage.py --xid 43 → app-bug-not-hardware. Software-induced fault — the fix is in the application (run it under compute-sanitizer), not an RMA. Pulling the card would waste healthy hardware and never fix the job.

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in skills/.curated/coreweave-gpu-node-forensics of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • ARD.md
  • PRD.md
  • eval-spec.yaml
  • references/cordon-rules.md
  • references/row-remap-decision.md
  • references/xid-triage-table.md
  • scripts/triage.py

Open the folder on GitHubat commit 23ea8d4

Compare with similar skills

Coreweave GPU Node Forensics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Coreweave GPU Node Forensics compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Coreweave GPU Node Forensics this skilljeremylongshore/tons-of-skills-marketplace2.8k—~3.2kAutomated safety check: PassMIT
Fla Triton To Gluonfla-org/flash-linear-attention5.8k—~4.2kAutomated safety check: PassMIT
DGX Spark Memory and Thermal Opswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
DGX Spark Training Gotchaswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
Physical AI Video Augmentation on OSMONVIDIA/skills3.5k—~4.7kAutomated safety check: NotesApache-2.0

Similar skills

  • Fla Triton To Gluon

    fla-org/flash-linear-attention

    Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…

    5.8k GitHub stars~4.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • OpenVLA-OFT Fine-Tuning

    Orchestra-Research/AI-Research-SKILLs

    Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

    13k GitHub starsUsed in 1 repo~3.7k tokens
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Analyzing Text With NLP

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to perform natural language processing and text analysis using the nlp-text-analyzer plugin.

    2.8k GitHub starsUsed in 1 repo~819 tokens
    Auto-check passed
  • Building Neural Networks

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill allows AI assistant to construct and configure neural network architectures using the neural-network-builder plugin.

    2.8k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Detecting Data Anomalies

    jeremylongshore/tons-of-skills-marketplace

    Process identify anomalies and outliers in datasets using machine learning algorithms.

    2.8k GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Explaining Machine Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill enables AI assistant to provide interpretability and explainability for machine learning models.

    2.8k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Optimizing Prompts

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill optimizes prompts for large language models (llms) to reduce token usage, lower costs, and improve performance.

    2.8k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed

Questions about Coreweave GPU Node Forensics

What does Coreweave GPU Node Forensics do?

Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently…. Coreweave GPU Node Forensics is an agent skill from jeremylongshore/tons-of-skills-marketplace. Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently kill a multi-day training run.

When should I use Coreweave GPU Node Forensics?

Coreweave GPU Node Forensics fits situations like: A GPU throws an Xid error; A node fell off the bus; A training run stalls; NCCL hangs on one rank.

How do I install Coreweave GPU Node Forensics in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a claude-code`. Or copy the skill folder (skills/.curated/coreweave-gpu-node-forensics in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/coreweave-gpu-node-forensics in your project. Claude Code loads it when a task matches its description.

How do I install Coreweave GPU Node Forensics in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a codex`. Or copy the skill folder (skills/.curated/coreweave-gpu-node-forensics in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/coreweave-gpu-node-forensics in your project. Codex loads it when a task matches its description.

Can I use Coreweave GPU Node Forensics in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-node-forensics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coreweave-gpu-node-forensics, .gemini/skills/coreweave-gpu-node-forensics, .github/skills/coreweave-gpu-node-forensics and .opencode/skills/coreweave-gpu-node-forensics in your project.

What does Coreweave GPU Node Forensics need to run?

Going by SKILL.md and its folder, Coreweave GPU Node Forensics needs Python for the scripts in its folder and the command-line tools its instructions call (python3, kubectl and jq). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Bash(python3:*), Bash(nvidia-smi -q:*), Bash(kubectl get:*), Bash(dmesg:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Coreweave GPU Node Forensics access the network?

SKILL.md names 1 domain. As links in the text: docs.nvidia.com. This is read from the text; nothing was executed.

Is Coreweave GPU Node Forensics safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Coreweave GPU Node Forensics use?

Coreweave GPU Node Forensics is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Coreweave GPU Node Forensics use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Coreweave GPU Node Forensics?

Skills that share tags, products or a category with Coreweave GPU Node Forensics: Fla Triton To Gluon (fla-org/flash-linear-attention, 5.8k stars), DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars), DGX Spark Training Gotchas (wshobson/agents, 40k stars) and Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Coreweave GPU Node Forensics?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,821 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 8, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.