Official agent skill

Gds Diag

by NVIDIA in NVIDIA/MagnumIO

A skill your agent uses when diagnosing NVIDIA GPUDirect Storage with this repository: choose and run the right gds-diag.py subcommand, interpret its output, and explain operator next steps without…

OfficialApache-2.0Auto-check passed

Install Gds Diag

skills CLI
$ npx skills add NVIDIA/MagnumIO --skill gds-diag -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/MagnumIO gds-diag --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/MagnumIO.git skills-src && mkdir -p .claude/skills && cp -r skills-src/gds-diag/skills/gds-diag .claude/skills/gds-diag && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gds-diag
GitHub stars
125
Token cost
~1.6k tokens
SKILL.md length
798 words
Files
5 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when diagnosing NVIDIA GPUDirect Storage with this repository: choose and run the right gds-diag.py subcommand, interpret its output, and explain operator next steps without…

  • Works in 7 steps: Identify the operator's intent and target. → Select the narrowest gds-diag.py… → Run the deterministic CLI from the… → …
  • Diagnosing NVIDIA GPUDirect Storage with this repository: choose and run the right gds-diag.py subcommand
  • SKILL.md covers Workflow, Command Routing, Interpretation Rules and Escalation
  • Calls python3

What it does

Gds Diag is an agent skill from NVIDIA/MagnumIO, published by the product's own GitHub organization. Use when diagnosing NVIDIA GPUDirect Storage with this repository: choose and run the right gds-diag.py subcommand, interpret its output, and explain operator next steps without duplicating the deterministic Python checks.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/command-routing.md`, `references/operator-next-steps.md` and `references/result-interpretation.md`).

It works with NVIDIA AI Platform and Python. The repository describes itself as: Magnum IO community repo. The licence is Apache-2.0.

When your agent uses it

  • Diagnosing NVIDIA GPUDirect Storage with this repository: choose and run the right gds-diag.py subcommand
  • Interpret its output
  • Explain operator next steps without duplicating the deterministic Python checks

Example prompts

  • “/gds-diag”

Requirements

  • Python 3
  • Docker

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Identify the operator's intent and target.
  2. Select the narrowest gds-diag.py subcommand that answers the question.
  3. Run the deterministic CLI from the repository root, or use
  4. If a runtime validation command is likely limited by the agent sandbox, ask
  5. Prefer verbose human output for operator-facing investigation. Prefer JSON
  6. Summarize the relevant WARN and FAIL findings first, then explain what they
  7. Give concrete next commands or documentation links only when they follow from

What it can do on your machine

Read from SKILL.md and the folder at commit 69b9d07. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gds Diag loads about 1.6k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 58 tokens; SKILL.md has 798 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/MagnumIO at commit 69b9d07, republished under its Apache-2.0 licence (© NVIDIA). 798 words, ~1,626 tokens.

Download SKILL.mdSave it as .claude/skills/gds-diag/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
gds-diag
description
Use when diagnosing NVIDIA GPUDirect Storage with this repository: choose and run the right gds-diag.py subcommand, interpret its output, and explain operator next steps without duplicating the deterministic Python checks.
SPDX-FileCopyrightText
Copyright (c) 2026, NVIDIA CORPORATION & AFFILIATES. All rights reserved.
SPDX-License-Identifier
CC-BY-4.0 AND Apache-2.0

GDS Diag

Use this skill to help an operator diagnose NVIDIA GPUDirect Storage (GDS) with the deterministic CLI in this repository.

The Python code is the source of truth for probing, parsing, support-matrix semantics, cufile.json validation, topology parsing, and output formatting. Do not reimplement those checks in this skill. Use the skill to understand the operator's goal, select the right subcommand, run it, and interpret its output.

Workflow

  1. Identify the operator's intent and target.
  2. Select the narrowest gds-diag.py subcommand that answers the question.
  3. Run the deterministic CLI from the repository root, or use scripts/gds-diag-wrapper from this skill directory when that is more reliable.
  4. If a runtime validation command is likely limited by the agent sandbox, ask for permission to rerun the same command outside the sandbox before treating the sandboxed output as host truth. See "Sandbox-limited NVIDIA runtime checks" below.
  5. Prefer verbose human output for operator-facing investigation. Prefer JSON output when you need stable structured data for follow-up reasoning.
  6. Summarize the relevant WARN and FAIL findings first, then explain what they mean for direct GDS, P2PDMA, RDMA, or compat mode.
  7. Give concrete next commands or documentation links only when they follow from the script output, local evidence, or current NVIDIA documentation.

Command Routing

Read references/command-routing.md when choosing a subcommand.

Common routes:

  • Broad or unclear GDS diagnosis, general bug-report collection, or "check this system/container/path": python3 gds-diag.py all PATH -v Use this as the default starting point when the operator has not already narrowed the problem. It detects host vs container context, runs the appropriate deterministic subcommands, and stops at the first unsuccessful return code. Use --json when collecting a structured report for automation or a bug.
  • Container GDS visibility, launch options, or comparing good/bad container runtime exposure: python3 gds-diag.py container-check -v Run this first from inside the target Docker or Enroot container before using post-install, mount-check, or support-matrix --live to diagnose that container. It distinguishes missing container devices, mounts, tools, and metadata from host-level GDS installation problems.
  • Host readiness before CUDA/GDS is fully installed: python3 gds-diag.py pre-install
  • Installed runtime validation: python3 gds-diag.py post-install -v
  • Path-specific compat-mode or storage-route diagnosis: python3 gds-diag.py mount-check PATH -v
  • Effective cuFile configuration audit from gdscheck -p, with file fallback and CUFILE_* environment overlays: python3 gds-diag.py config-audit --profile PROFILE
  • Proposed or copied cufile.json/JSONC audit: python3 gds-diag.py config-audit --config PATH --ignore-env -v
  • Filesystem support reference: python3 gds-diag.py support-matrix
  • Installed/runtime support view: python3 gds-diag.py support-matrix --live

Interpretation Rules

Read references/result-interpretation.md before explaining non-trivial output.

Important boundaries:

  • Do not treat compat mode as direct GDS. It is a CPU bounce-buffer fallback.
  • Do not claim a filesystem supports P2PDMA/C2C just because a global JSON setting is enabled. Route support is limited by the Python support matrix.
  • NFS's direct GDS path is NFSoRDMA, not the NVMe-style nvfs path — but nvidia-fs (nvfs) still has to be loaded to activate it, so NFS is Native-applicable, same as Lustre and BeeGFS. Don't describe NFS as supporting native GDS "via NVMe" — it's a different mechanism — but also don't claim NFS has no relationship to nvidia-fs/nvfs at all.
  • For GPFS and WekaFS, focus on userspace RDMA via DmaBuf or nvidia_peermem, not NVMe-style nvfs.
  • For release-specific behavior, rely on the tool's version-aware output and current NVIDIA documentation.
Show full SKILL.md (281 more words)Show less

Escalation

Some commands may need sudo or host-specific access to gather complete evidence. Ask before using privileged commands unless the user already asked for that level of probing.

Sandbox-limited NVIDIA runtime checks

Agent execution sandboxes may hide /dev/nvidia* device nodes even when the real host has a working NVIDIA driver. This can make nvidia-smi, gdscheck, GPU topology checks, and post-install runtime validation fail inside the agent while succeeding in the operator's normal terminal.

For commands that validate installed runtime state, especially:

  • python3 gds-diag.py post-install -v
  • python3 gds-diag.py mount-check PATH -v
  • python3 gds-diag.py all PATH -v
  • python3 gds-diag.py support-matrix --live

use this flow:

  1. Run the selected command normally first.

  2. If it fails because nvidia-smi cannot communicate with the NVIDIA driver, the post-install prerequisite gate says runtime GPU validation is unavailable, gdscheck cannot access the runtime, or /dev/nvidia* appears missing from the agent environment while lspci, modinfo nvidia, /proc/driver/nvidia/version, or /proc/devices indicate the driver/GPU exists, do not conclude that the host driver is broken.

  3. Request permission to rerun the exact same gds-diag.py command outside the sandbox using the agent's escalation mechanism. In Codex, run the command with sandbox_permissions="require_escalated" and a justification such as:

    text
    Allow running gds-diag post-install outside the sandbox so it can access /dev/nvidia* and validate the live NVIDIA/GDS runtime?
  4. Treat the outside-sandbox result as authoritative for host status. If the user declines escalation, say that the result is limited by agent sandbox visibility and ask the operator to run the same command in a normal host terminal.

Do not change the deterministic CLI result text in the skill. The skill's job is to choose the right execution environment and explain when sandbox visibility limits the evidence.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in gds-diag/skills/gds-diag of NVIDIA/MagnumIO.

  • SKILL.md
  • references/command-routing.md
  • references/operator-next-steps.md
  • references/result-interpretation.md
  • scripts/gds-diag-wrapper

Open the folder on GitHubat commit 69b9d07

Compare with similar skills

Gds Diag next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gds Diag compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gds Diag this skillNVIDIA/MagnumIO125—~1.6kAutomated safety check: PassApache-2.0
Refactor OpCVCUDA/CV-CUDA2.7k—~1.5kAutomated safety check: PassCustom licence
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Nsight Graphics AnalyzerLuna5ama/Alpha-Piscium156—~4.7kAutomated safety check: PassGPL-3.0
Optimize OpCVCUDA/CV-CUDA2.7k—~834Automated safety check: PassCustom licence
Deep Researcher ResearchNVIDIA-AI-Blueprints/deep-researcher-agent885—~4.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Refactor Op

    CVCUDA/CV-CUDA

    Find and safely apply per-operator refactoring / redundancy-reduction opportunities in a CV-CUDA operator (near-duplicate Tensor/VarShape kernels, reinvented shared utilities, dead code).

    2.7k GitHub stars~1.5k tokensUpdated 23 days ago
    DevelopmentAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Nsight Graphics Analyzer

    Luna5ama/Alpha-Piscium

    Drive NVIDIA Nsight Graphics 2026.1+ from the command line for GPU performance analysis, frame capture, frame trace inspection, draw-call inspection, NVTX/D3DPERF stage timing, replay metadata…

    156 GitHub stars~4.7k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Optimize Op

    CVCUDA/CV-CUDA

    Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

    2.7k GitHub stars~834 tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Deep Researcher Research

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses when asked to run deep research or Deep Researcher Agent research through a reachable NVIDIA Deep Researcher Agent Blueprint backend.

    885 GitHub stars~4.4k tokensUpdated yesterday
    Research & ScienceAuto-check: notes
  • Object Separation

    Barty-Bart/motion-graphics

    Separate a person, product or hand from the background in a video, on the user's own computer with SAM 2.

    518 GitHub stars~2.4k tokensUpdated yesterday
    Media & CreativeAuto-check passed

Questions about Gds Diag

What does Gds Diag do?

A skill your agent uses when diagnosing NVIDIA GPUDirect Storage with this repository: choose and run the right gds-diag.py subcommand, interpret its output, and explain operator next steps without…. Gds Diag is an agent skill from NVIDIA/MagnumIO, published by the product's own GitHub organization.py subcommand, interpret its output, and explain operator next steps without duplicating the deterministic Python checks.

When should I use Gds Diag?

Gds Diag fits situations like: diagnosing NVIDIA GPUDirect Storage with this repository: choose and run the right gds-diag.py subcommand; interpret its output; explain operator next steps without duplicating the deterministic Python checks.

How do I install Gds Diag in Claude Code?

Run `npx skills add NVIDIA/MagnumIO --skill gds-diag -a claude-code`. Or copy the skill folder (gds-diag/skills/gds-diag in NVIDIA/MagnumIO) into .claude/skills/gds-diag in your project. Claude Code loads it when a task matches its description.

How do I install Gds Diag in Codex?

Run `npx skills add NVIDIA/MagnumIO --skill gds-diag -a codex`. Or copy the skill folder (gds-diag/skills/gds-diag in NVIDIA/MagnumIO) into .agents/skills/gds-diag in your project. Codex loads it when a task matches its description.

Can I use Gds Diag in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/MagnumIO --skill gds-diag -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gds-diag, .gemini/skills/gds-diag, .github/skills/gds-diag and .opencode/skills/gds-diag in your project.

What does Gds Diag need to run?

Going by SKILL.md and its folder, Gds Diag needs the command-line tools its instructions call (python3). Our summary lists: Python 3; Docker.

Does Gds Diag access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gds Diag safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gds Diag use?

Gds Diag is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gds Diag use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.

What are the alternatives to Gds Diag?

Skills that share tags, products or a category with Gds Diag: Refactor Op (CVCUDA/CV-CUDA, 2.7k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Nsight Graphics Analyzer (Luna5ama/Alpha-Piscium, 156 stars) and Optimize Op (CVCUDA/CV-CUDA, 2.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gds Diag?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/MagnumIO, which has 125 GitHub stars. The repository was last updated on October 8, 2026.

Source: NVIDIA/MagnumIO on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.