Official agent skill

Dynamo Interconnect Check

by NVIDIA in NVIDIA/skills

Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Dynamo Interconnect Check

skills CLI
$ npx skills add NVIDIA/skills --skill dynamo-interconnect-check -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills dynamo-interconnect-check --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dynamo-interconnect-check .claude/skills/dynamo-interconnect-check && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dynamo-interconnect-check
GitHub stars
3.5k
Used in
1 other repo
Token cost
~1.6k tokens
SKILL.md length
657 words
Files
7 (incl. scripts, references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink.

  • Works in 3 steps: Check Transport Env Vars On The Recipe → Check Node Capabilities → Validate NIXL Reachability
  • Tasks that involve GPU and accelerator computing
  • SKILL.md covers Purpose, Prerequisites, When To Use and Instructions, plus 7 more sections
  • Runs Python scripts from its folder; calls python3 and kubectl

What it does

Dynamo Interconnect Check is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. Use after recipe-runner brings a deployment up (especially disagg/multi-node) to confirm the KV transport is correct; use troubleshoot for diagnosing already-failed pods.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/interconnect-env-vars.md`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing and Deployment. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve GPU and accelerator computing
  • Tasks that involve Deployment

Example prompts

  • “/dynamo-interconnect-check”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Check Transport Env Vars On The Recipe
  2. Check Node Capabilities
  3. Validate NIXL Reachability

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dynamo Interconnect Check loads about 1.6k tokens when it runs, and up to ~2.5k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 657 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 657 words, ~1,599 tokens.

Download SKILL.mdSave it as .claude/skills/dynamo-interconnect-check/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
dynamo-interconnect-check
description
Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. Use after recipe-runner brings a deployment up (especially disagg/multi-node) to confirm the KV transport is correct; use troubleshoot for diagnosing already-failed pods.
license
Apache-2.0
metadata.author
Dan Gil <dagil@nvidia.com>
metadata.tags
dynamo, nixl, rdma, disagg, validation

Dynamo Interconnect Check

<!--
SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
SPDX-License-Identifier: CC-BY-4.0
-->

Purpose

Confirm that the transport disaggregated serving depends on actually works. A deployment can pass an endpoint smoke test while disagg is silently wrong: if NIXL/UCX cannot reach the peer worker over RDMA or NVLink, KV transfer falls back to a slow or broken path. Catch that with read-only checks before trusting a disagg deployment or its benchmark numbers.

This skill is read-only. It never mutates the cluster and never prints secrets.

Prerequisites

  • Python 3.10+ on the operator machine.
  • kubectl exec access to a worker pod in the target Dynamo deployment.
  • Read access to the recipe directory (recipes/<model>/<framework>/<mode>).
  • For node-capability checks: tools like ibstat, nvidia-smi, lsmod available in the worker pod image (missing tools are reported as skipped, not failures).

When To Use

  • After dynamo-recipe-runner deploys a disagg or multi-node recipe.
  • Before reporting disagg throughput/latency, so numbers reflect the real transport.
  • When agg works but disagg is slow, hangs, or returns wrong output and you suspect the fabric rather than the model.

For diagnosing pods that are already crashing or unschedulable, use dynamo-troubleshoot first.

Instructions

1. Check Transport Env Vars On The Recipe
bash
python3 scripts/check_interconnect.py env recipes/<model>/<framework>/<mode>

Reports which NIXL/UCX/NCCL transport variables are set and flags disagg-critical ones (e.g. UCX_TLS, UCX_NET_DEVICES, NCCL_IB_HCA) that are absent. Missing here is only a warning — they may be baked into the image — so confirm with the node and NIXL checks. See references/interconnect-env-vars.md for what each variable does.

2. Check Node Capabilities

Locally on a GPU node, or inside a running worker pod:

bash
python3 scripts/check_interconnect.py node \
  --namespace "${NAMESPACE}" --pod <worker-pod>

Probes (read-only) for: InfiniBand devices and Active links, GPUDirect RDMA (nvidia_peermem), GDRCopy, and NVLink in the GPU topology. Missing tools are reported as skipped, not failures.

3. Validate NIXL Reachability
bash
python3 scripts/check_interconnect.py nixl \
  --namespace "${NAMESPACE}" --pod <worker-pod>

Looks for NIXL test tooling in the pod and surfaces the exact next step to run a pairwise prefill↔decode transfer test. A full cross-pod transfer test requires two scheduled GPU pods on the fabric.

Available Scripts

ScriptPurposeArguments
scripts/check_interconnect.py envInspect NIXL/UCX/NCCL env vars on a recipepositional recipe path
scripts/check_interconnect.py nodeProbe InfiniBand, GPUDirect RDMA, GDRCopy, NVLink on a node or pod--namespace, --pod
scripts/check_interconnect.py nixlSurface NIXL transfer-test readiness for a pod--namespace, --pod

Invoke via the agentskills.io run_script() protocol:

python
run_script("scripts/check_interconnect.py", args=["env", "recipes/qwen3-coder-480b/sglang/disagg"])
run_script("scripts/check_interconnect.py", args=["node", "--namespace", "dynamo-demo", "--pod", "qwen-worker-0"])

Examples

Verify a disagg recipe's transport env shape before deploy:

bash
python3 scripts/check_interconnect.py env recipes/qwen3-coder-480b/sglang/disagg

After deploy, validate a worker pod's fabric:

bash
python3 scripts/check_interconnect.py node \
  --namespace dynamo-demo --pod qwen-worker-0
python3 scripts/check_interconnect.py nixl \
  --namespace dynamo-demo --pod qwen-worker-0

Equivalent through the agent protocol:

python
run_script("scripts/check_interconnect.py", args=["nixl", "--namespace", "dynamo-demo", "--pod", "qwen-worker-0"])
Show full SKILL.md (275 more words)Show less

Output Contract

Each check returns ok / warn / fail / skipped with a one-line detail, plus a rolled-up verdict on disagg transport readiness. Report:

  • transport env vars present vs. disagg-critical ones missing
  • RDMA / GPUDirect / NVLink capability status
  • whether NIXL reachability was validated, and the next command if not
  • a clear statement of whether disagg can be trusted, or what to fix first

Limitations

  • Read-only fabric probe; does not run a full pairwise NIXL transfer (requires two scheduled GPU pods and the in-pod NIXL test tools).
  • skipped results for missing tools (ibstat, nvidia-smi, lsmod) are inconclusive, not a pass.
  • Env-var check inspects the recipe text; values injected at runtime via initContainers or operator-applied envs are not detected.
  • Single-node agg deployments do not exercise the transport — this skill is for disagg / multi-node validation.

Troubleshooting

SymptomLikely causeNext step
env reports all critical vars missingVars baked into image or injected by operatorRun the node check inside the worker pod to verify actual env
node reports no Active IB linkFabric down or HCA not provisioned to the nodeContact cluster admin; verify kubectl describe node shows nvidia.com/gpu and IB labels
nvidia_peermem missingGPUDirect RDMA module not loadedAsk cluster admin to load nvidia-peermem; without it, NIXL falls back to staged copies
nixl finds no test toolsWorker image lacks NIXL test harnessUse a NIXL-enabled image or run the standalone transfer test from a debug pod

Benchmark

See BENCHMARK.md for the NVCARPS-EVAL performance report (auto-generated by the NVSkills CI pipeline). To refresh, re-run /nvskills-ci on an upstream PR touching this skill.

References

  • references/interconnect-env-vars.md — NIXL/UCX/NCCL env var catalog and IB capability checklist.
  • Use scripts/check_interconnect.py for all read-only checks.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in skills/dynamo-interconnect-check of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • references/interconnect-env-vars.md
  • scripts/check_interconnect.py
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 0e0d506

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NVIDIA/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Dynamo Interconnect Check next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dynamo Interconnect Check compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dynamo Interconnect Check this skillNVIDIA/skills3.5k1 repos~1.6kAutomated safety check: PassApache-2.0
Cv DeployLMIXR/CV_Deployment_skill126—~547Automated safety check: PassNone
TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs13k5 repos~1.3kAutomated safety check: PassMIT
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
DGX Spark Memory and Thermal Opswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
Setup Workshopbrevdev/workshop-build-an-agent143—~2.3kAutomated safety check: NotesApache-2.0

Similar skills

  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    126 GitHub stars~547 tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • TensorRT-LLM Inference

    Orchestra-Research/AI-Research-SKILLs

    Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

    13k GitHub starsUsed in 5 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop

    brevdev/workshop-build-an-agent

    This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a.

    143 GitHub stars~2.3k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Nemotron Customizer Airgap

    NVIDIA-NeMo/Nemotron

    Prepare, validate, build, and use Nemotron Customizer airgap image bundles for offline clusters.

    2.1k GitHub stars~1.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Dynamo Interconnect Check

What does Dynamo Interconnect Check do?

Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. Dynamo Interconnect Check is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink.

When should I use Dynamo Interconnect Check?

Dynamo Interconnect Check fits situations like: tasks that involve GPU and accelerator computing; tasks that involve Deployment.

How do I install Dynamo Interconnect Check in Claude Code?

Run `npx skills add NVIDIA/skills --skill dynamo-interconnect-check -a claude-code`. Or copy the skill folder (skills/dynamo-interconnect-check in NVIDIA/skills) into .claude/skills/dynamo-interconnect-check in your project. Claude Code loads it when a task matches its description.

How do I install Dynamo Interconnect Check in Codex?

Run `npx skills add NVIDIA/skills --skill dynamo-interconnect-check -a codex`. Or copy the skill folder (skills/dynamo-interconnect-check in NVIDIA/skills) into .agents/skills/dynamo-interconnect-check in your project. Codex loads it when a task matches its description.

Can I use Dynamo Interconnect Check in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill dynamo-interconnect-check -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dynamo-interconnect-check, .gemini/skills/dynamo-interconnect-check, .github/skills/dynamo-interconnect-check and .opencode/skills/dynamo-interconnect-check in your project.

What does Dynamo Interconnect Check need to run?

Going by SKILL.md and its folder, Dynamo Interconnect Check needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and kubectl). Our summary lists: Python 3.

Does Dynamo Interconnect Check access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dynamo Interconnect Check safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Dynamo Interconnect Check use?

Dynamo Interconnect Check is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dynamo Interconnect Check use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 895 tokens, read only when the agent opens those files.

What are the alternatives to Dynamo Interconnect Check?

Skills that share tags, products or a category with Dynamo Interconnect Check: Cv Deploy (LMIXR/CV_Deployment_skill, 126 stars), TensorRT-LLM Inference (Orchestra-Research/AI-Research-SKILLs, 13k stars), Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars) and DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dynamo Interconnect Check?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.