Official agent skill

Hyperpod Nccl

by awslabs in awslabs/agent-plugins

Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Hyperpod Nccl

skills CLI
$ npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install awslabs/agent-plugins hyperpod-nccl --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-nccl .claude/skills/hyperpod-nccl && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hyperpod-nccl
GitHub stars
915
Token cost
~3.4k tokens
SKILL.md length
984 words
Files
6 (incl. scripts, references)
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…

  • Works in 2 steps: Authenticate kubectl (EKS) → Run the diagnostic
  • Tasks that involve GPU and accelerator computing
  • SKILL.md covers Workflow, Step 1: Authenticate kubectl…, Step 2: Run the diagnostic and Remediation index, plus 8 more sections
  • Runs Shell scripts from its folder; calls aws, bash and kubectl

What it does

Hyperpod Nccl is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures, EFA TCP fallback, /dev/shm or memlock issues, NCCL version mismatch across pods, container OOM / exit-137 / OOMKilled, GPU OOM (CUDA out of memory), CrashLoopBackOff / Pending pods, MASTERADDR DNS, NetworkPolicy blocking. Not for single-node hardware faults (→ hyperpod-node-debugger § G) or cluster-creation EFA / SSM…

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/debugging-guide.md`, `references/error-patterns-quick-ref.md` and `references/operations.md`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with Amazon SageMaker, Amazon Web Services and CUDA. The repository describes itself as: Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve GPU and accelerator computing

Example prompts

  • “/hyperpod-nccl”

Requirements

  • A Bash shell

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Authenticate kubectl (EKS)
  2. Run the diagnostic

What it can do on your machine

Read from SKILL.md and the folder at commit da51970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • aws
    • bash
    • kubectl
    • yum
    • apt

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use aws and kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hyperpod Nccl loads about 3.4k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 145 tokens; SKILL.md has 984 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~145
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~25k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from awslabs/agent-plugins at commit da51970, republished under its Apache-2.0 licence (© awslabs). 984 words, ~3,420 tokens.

Download SKILL.mdSave it as .claude/skills/hyperpod-nccl/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
hyperpod-nccl
description
Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures, EFA TCP fallback, /dev/shm or memlock issues, NCCL version mismatch across pods, container OOM / exit-137 / OOMKilled, GPU OOM (CUDA out of memory), CrashLoopBackOff / Pending pods, MASTER_ADDR DNS, NetworkPolicy blocking. Not for single-node hardware faults (→ hyperpod-node-debugger § G) or cluster-creation EFA / SSM failures (→ hyperpod-cluster-debugger § A / § F).
metadata.version
0.0.1

HyperPod NCCL Debugger

Operating policy. Run read-only diagnostics yourself. Never run a command that changes cluster, node, or workload state — present each one as a Suggested command (run this yourself) block and wait for the customer. Destructive order: investigate → reboot → replace (replace destroys root + secondary volumes; not supported on Slurm controller nodes). Never discard training state on speculation.

Diagnose NCCL failures on SageMaker HyperPod (EKS and Slurm). scripts/nccl-diagnose.sh reads state via AWS APIs, kubectl, and SSM, then prints each issue as [FAIL] ... → references/<file>.md § <section>. Read-only.

Signal sourcing: list-cluster-events carries infrastructure-level state only (lifecycle, bootstrap, EFA health check, capacity, replacement, reboot, AMI rollback). It does not carry NCCL timeouts, GPU XID/ECC, or per-pod training signals — those come from pod logs, CloudWatch training streams, on-node SSM probes, and NCCL env audit. "No events" on a training-time NCCL issue is expected, not a clean bill of health.


Workflow

  1. Collect cluster name, region, namespace/job (EKS), exact NCCL error string.
  2. Run the diagnostic (always — the output drives everything else).
  3. For every [FAIL] line, Read the referenced section.
  4. Present finding, root cause, and the Suggested-command block with concrete values (instance IDs, SG IDs, namespaces) filled in from the script output. Wait for customer approval.
  5. Re-run the diagnostic to confirm.

If a finding has no matching section, report it as a bug — do not invent a fix.

Step 1: Authenticate kubectl (EKS)

bash
EKS_ARN=$(aws sagemaker describe-cluster --cluster-name <HYPERPOD-NAME> --region <REGION> \
  --query 'Orchestrator.Eks.ClusterArn' --output text)
EKS_NAME=$(echo "$EKS_ARN" | awk -F'/' '{print $NF}')
aws eks update-kubeconfig --name "$EKS_NAME" --region <REGION>
kubectl get nodes

Step 2: Run the diagnostic

bash
# Basic:
bash scripts/nccl-diagnose.sh --cluster <HYPERPOD-NAME> --region <REGION>

# Scope to an EKS job/namespace:
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --namespace <NS> --job <JOB>

# Force orchestrator:
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --orchestrator slurm

# Larger hardware sample (default 3):
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --sample-nodes 10

# Specific node only:
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --node i-0abc123def456

Tags: [PASS] · [FAIL] (counted in Issues Found, has reference pointer) · [WARN] · [INFO]. Priorities: P0 blocks training · P1 degraded · P2 informational.


Remediation index

Each [FAIL] line in the script already points directly at the right section. This table is a lookup for manual triage.

FindingSection
SG missing inbound/outbound self-referenceoperations.md § 8
Blocking NetworkPolicy / allow-all missingoperations.md § 8
Slurm node DOWN / DRAINING / RemoveIPCoperations.md § 7
GPU XID / SYSTEM_ERROR / hardware faulthyperpod-node-debugger § F / § G
GPU row-remap / DCGM Fail / silent NaNshyperpod-node-debugger § G.1.a/b
NCCL timeout / rendezvous / stragglerdebugging-guide.md § 1
EFA configuration / not useddebugging-guide.md § 6
EFA TCP fallback (NET/OFI Using TCP)debugging-guide.md § 13
NCCL version mismatch across podsdebugging-guide.md § 10
Container OOM (pod killed, exit 137)debugging-guide.md § 4
GPU OOM (CUDA out of memory)debugging-guide.md § 11
RDMA memlock / /dev/shm too smalldebugging-guide.md § 17
MASTER_ADDR DNS / headless Servicedebugging-guide.md § 12
NVLS / PXN / topology tuningdebugging-guide.md § 19
Any NCCL / EFA / rendezvous log patternerror-patterns-quick-ref.md
Performance / nccl-tests / bandwidthperformance-testing.md

Prerequisites

  • aws CLI v2.13+ authenticated (aws sts get-caller-identity)
  • jq, python3, bash 4.2+
  • unbuffer (from the expect package: yum install expect / apt install expect)
  • kubectl authenticated to the EKS cluster (K8s checks skipped if absent)
  • session-manager-plugin for on-node hardware checks

Defaults

  • Region — required: pass --region or set $AWS_DEFAULT_REGION.
  • Orchestrator — auto-detected; override with --orchestrator eks|slurm.
  • Namespace / job (EKS) — all namespaces; scope with --namespace <NS> --job <JOB>.
  • Hardware sampling — 3 nodes over SSM (capped at 50). --node <ID> for a specific node. Node probes run serially (180 s per node): --sample-nodes 10 can take ~30 min.
  • CloudWatch window — last 2 hours.
  • Colors — auto-disabled on non-TTY or TERM=dumb.

Error handling

FailureScriptTell the customer
aws sts get-caller-identity failsExit 1 with the AWS error"Fix AWS credentials and rerun."
describe-cluster AccessDeniedWarn, add Missing IAM for sagemaker:DescribeCluster"Grant sagemaker:DescribeCluster (operations.md § 2)."
Cluster not foundExit 1 after listing region's clusters"Confirm HyperPod cluster name and region."
kubectl absent / unauthenticatedWarn, skip K8s checks"aws eks update-kubeconfig --name <EKS> --region <R>."
SSM plugin absentWarn, skip on-node hardware checks"Install session-manager-plugin."
SSM times out (180s)Partial output, mark node unreachable"Rerun with --node <ID> --sample-nodes 1; check SSM agent on the node."
CloudWatch log group not foundSkip CloudWatch scan"Enable CloudWatch on the cluster (operations.md § 4)."
Cluster events API throttledWarn, continue with partial data"Rerun later — script is idempotent."

Exit codes: 0 diagnostic complete · 1 fatal prerequisite missing or cluster unreachable.

Show full SKILL.md (357 more words)Show less

IAM permissions

Full policy + RBAC in operations.md § 2. SSM on HyperPod uses start-session against sagemaker-cluster:<cluster-id>_<group>-<iid> targets — grant ssm:StartSession / ssm:TerminateSession, not ssm:SendCommand.

Scale strategy

ScopeMethodCoverage
All nodessagemaker:ListClusterNodes (paginated)100% nodes
All K8s objectskubectl100% pods/nodes/policies
HardwareSSM --sample-nodes N (default 3)Sampled
Node logsCloudWatch100% nodes

Large clusters: the PyTorch NCCL backend defaults to a 10-minute collective-op timeout (per the PyTorch distributed docs). Large clusters routinely exceed that on first rendezvous; raise it via torch.distributed.init_process_group(timeout=timedelta(seconds=<N>)). HyperPod support has also observed NCCL topology-graph-search hangs on 256+ node clusters when memlock is unlimited; using a large fixed memlock (e.g. 8388608) in pod securityContext or /etc/security/limits.conf has cleared these in field cases. This memlock pattern is a field observation, not AWS- or NCCL-documented behavior.

For FSDP, DeepSpeed, or Megatron-LM tuning: debugging-guide.md § 18.

Skill delegation

NeedUse
Cluster creation / deployment failureshyperpod-cluster-debugger (§ A / B / C / H + --validate)
Post-deployment cluster-wide managementhyperpod-cluster-debugger
Per-node issues (disk, lifecycle, hardware)hyperpod-node-debugger
Trainium/Inferentia collective-comm (AWS Neuron Collectives, not NCCL)hyperpod-node-debugger § G.2
Shell on nodeshyperpod-ssm
Version comparison across nodeshyperpod-version-checker
Diagnostic bundle for AWS Supporthyperpod-issue-report
MFU / performance degradationhyperpod-mfu-debugger

Escalate to AWS Support

Escalate when:

  1. All SG rules correct, EFA verified on-node, but NCCL still times out.
  2. Hardware checks pass on all nodes but AllReduce still hangs.
  3. Issues Found: 0 but training still fails.
  4. GPU XID errors persist after node replacement.
  5. Collective-op timeout raised and memlock workaround applied but large-cluster rendezvous still hangs.
Before opening the case
bash
# 1. Cluster identity + status
aws sagemaker describe-cluster --cluster-name <C> --region <R>

# 2. Full NCCL diagnostic (sample more nodes for escalation)
bash scripts/nccl-diagnose.sh --cluster <C> --region <R> --sample-nodes 10 > nccl-diag.txt

# 3. Per-node log/config bundle to S3 (delegates to hyperpod-issue-report)
#    See skills/hyperpod-issue-report/SKILL.md for the exact invocation.
Include in the case
  • Cluster name + ARN and AWS region
  • Orchestrator (EKS or Slurm) and EKS cluster name / Slurm controller node
  • Timestamp window (UTC start / end) of the failure
  • Exact NCCL / libfabric error strings (copy verbatim from pod logs or journalctl)
  • Affected instance IDs / node names / pod names / namespace / job name
  • nccl-diag.txt from step 2 above
  • S3 URI of the hyperpod-issue-report bundle from step 3
  • NCCL env vars in effect (printenv | grep -E '^NCCL|^FI_|^TORCH_' from one pod)

References

© awslabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in plugins/sagemaker-ai/skills/hyperpod-nccl of awslabs/agent-plugins.

  • SKILL.md
  • references/debugging-guide.md
  • references/error-patterns-quick-ref.md
  • references/operations.md
  • references/performance-testing.md
  • scripts/nccl-diagnose.sh

Open the folder on GitHubat commit da51970

Compare with similar skills

Hyperpod Nccl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hyperpod Nccl compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hyperpod Nccl this skillawslabs/agent-plugins915—~3.4kAutomated safety check: PassApache-2.0
Aiml GPU Training Cluster Investigationaws/tools-for-devops-agent100—~5.4kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Python Environment Setup for SageMakerhuggingface/skills11k2 repos~1.7kAutomated safety check: PassApache-2.0
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence

Similar skills

  • Aiml GPU Training Cluster Investigation

    aws/tools-for-devops-agent

    Official

    A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.

    100 GitHub stars~5.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed

More from awslabs/agent-plugins

All 33 skills in this repo
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    915 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    915 GitHub starsUsed in 2 repos~3.5k tokens
    Auto-check passed
  • Finetuning Technique

    awslabs/agent-plugins

    Official

    Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes.

    915 GitHub starsUsed in 1 repo~604 tokens
    Auto-check passed
  • Hyperpod Issue Report

    awslabs/agent-plugins

    Official

    Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases.

    915 GitHub starsUsed in 1 repo~890 tokens
    Auto-check passed
  • Hyperpod Performance Debugger

    awslabs/agent-plugins

    Official

    Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

    915 GitHub starsUsed in 1 repo~4.1k tokens
    Auto-check passed
  • Hyperpod Ssm

    awslabs/agent-plugins

    Official

    Remote command execution and file transfer on SageMaker HyperPod cluster nodes via AWS Systems Manager (SSM).

    915 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check: notes

Questions about Hyperpod Nccl

What does Hyperpod Nccl do?

Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…. Hyperpod Nccl is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures, EFA TCP fallback, /dev/shm or memlock issues, NCCL version mismatch across pods, container OOM / exit-137 / OOMKilled, GPU OOM (CUDA out of memory), CrashLoopBackOff / Pending pods, MASTERADDR DNS, NetworkPolicy blocking.

When should I use Hyperpod Nccl?

Hyperpod Nccl fits situations like: tasks that involve GPU and accelerator computing.

How do I install Hyperpod Nccl in Claude Code?

Run `npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a claude-code`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-nccl in awslabs/agent-plugins) into .claude/skills/hyperpod-nccl in your project. Claude Code loads it when a task matches its description.

How do I install Hyperpod Nccl in Codex?

Run `npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a codex`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-nccl in awslabs/agent-plugins) into .agents/skills/hyperpod-nccl in your project. Codex loads it when a task matches its description.

Can I use Hyperpod Nccl in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hyperpod-nccl, .gemini/skills/hyperpod-nccl, .github/skills/hyperpod-nccl and .opencode/skills/hyperpod-nccl in your project.

What does Hyperpod Nccl need to run?

Going by SKILL.md and its folder, Hyperpod Nccl needs a shell for the scripts in its folder and the command-line tools its instructions call (aws, bash, kubectl, yum and apt). Our summary lists: A Bash shell.

Does Hyperpod Nccl access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Hyperpod Nccl safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Hyperpod Nccl use?

Hyperpod Nccl is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hyperpod Nccl use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 22k tokens, read only when the agent opens those files.

What are the alternatives to Hyperpod Nccl?

Skills that share tags, products or a category with Hyperpod Nccl: Aiml GPU Training Cluster Investigation (aws/tools-for-devops-agent, 100 stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars) and Python Environment Setup for SageMaker (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hyperpod Nccl?

awslabs (a GitHub organization, an official publisher) maintains it in awslabs/agent-plugins, which has 915 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 5, 2026.

Source: awslabs/agent-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.