Official agent skill

Hyperpod Node Debugger

by awslabs in awslabs/agent-plugins

Diagnose and remediate per-node issues on a HyperPod cluster (EKS or Slurm) — a specific node is unhealthy, unresponsive, stuck, or needs replacing.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Hyperpod Node Debugger

skills CLI
$ npx skills add awslabs/agent-plugins --skill hyperpod-node-debugger -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install awslabs/agent-plugins hyperpod-node-debugger --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-node-debugger .claude/skills/hyperpod-node-debugger && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hyperpod-node-debugger
GitHub stars
915
Token cost
~5.2k tokens
SKILL.md length
1,619 words
Files
7 (incl. scripts, references)
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnose and remediate per-node issues on a HyperPod cluster (EKS or Slurm) — a specific node is unhealthy, unresponsive, stuck, or needs replacing.

  • Works in 2 steps: Triage → Match signal → section
  • Tasks that involve GPU and accelerator computing
  • SKILL.md covers Workflow, Step 1: Triage, Step 2: Match signal → section and A: EFA / Security Group, plus 20 more sections
  • Runs Shell scripts from its folder; calls bash, aws and kubectl

What it does

Hyperpod Node Debugger is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose and remediate per-node issues on a HyperPod cluster (EKS or Slurm) — a specific node is unhealthy, unresponsive, stuck, or needs replacing. Covers on-node EFA, GPU / accelerator hardware (XID, ECC, NVLink, row-remap, DCGM), Slurm node down/drained, disk and memory pressure, per-node lifecycle-script failures, SSM agent, container runtime, kernel panics, pod networking. Read-only. Not for cluster-wide provisioning (→ hyperpod-cluster-debugger), NCCL (→ hyperpod-nccl), or MFU (→ hyperpod-mfu-debugger).

Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/node-diagnostics-detail.md`, `references/node-issue-catalog.md` and `scripts/check-efa-sg.sh`).

It sits in DevOps & Cloud, covering GPU and accelerator computing. The repository describes itself as: Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve GPU and accelerator computing

Example prompts

  • “/hyperpod-node-debugger”

Requirements

  • A Bash shell
  • Docker

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Triage
  2. Match signal → section

What it can do on your machine

Read from SKILL.md and the folder at commit da51970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • aws
    • kubectl
    • yum
    • apt

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use aws and kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hyperpod Node Debugger loads about 5.2k tokens when it runs, and up to ~24k if it reads all its reference files. Until then it costs about 134 tokens; SKILL.md has 1,619 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~5.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~24k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from awslabs/agent-plugins at commit da51970, republished under its Apache-2.0 licence (© awslabs). 1,619 words, ~5,201 tokens.

Download SKILL.mdSave it as .claude/skills/hyperpod-node-debugger/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
hyperpod-node-debugger
description
Diagnose and remediate per-node issues on a HyperPod cluster (EKS or Slurm) — a specific node is unhealthy, unresponsive, stuck, or needs replacing. Covers on-node EFA, GPU / accelerator hardware (XID, ECC, NVLink, row-remap, DCGM), Slurm node down/drained, disk and memory pressure, per-node lifecycle-script failures, SSM agent, container runtime, kernel panics, pod networking. Read-only. Not for cluster-wide provisioning (→ hyperpod-cluster-debugger), NCCL (→ hyperpod-nccl), or MFU (→ hyperpod-mfu-debugger).
metadata.version
0.0.1

HyperPod Node Debugger

Operating policy. Run read-only diagnostics yourself. Never run a command that changes cluster, node, or workload state — present each one as a Suggested command (run this yourself) block and wait for the customer. Destructive order: investigate → reboot → replace (replace destroys root + secondary volumes; not supported on Slurm controller nodes). Never discard training state, logs, or caches on speculation.

IaC note (always include with mutation commands). When you suggest any command that changes cluster, VPC, SG, subnet, or EKS configuration (e.g. authorize-security-group-*, modify-vpc-attribute, update-cluster, kubectl label/cordon/drain, create namespace, set env daemonset), ask the customer first whether the cluster / VPC / SG is managed by Infrastructure-as-Code (CloudFormation, CDK, Terraform, Pulumi). If yes, tell them: "Apply this change in your IaC source first, then deploy through the pipeline — running the command directly will drift from your template and the next stack update may overwrite it." If they need to fix the issue immediately and the IaC change will follow, flag the drift explicitly so they remember to reconcile.

Read-only triage. scripts/triage-cluster.sh (and helpers check-efa-sg.sh, check-node-reachability.sh, check-vpc-config.sh) read state and print each issue as [FAIL] ... → references/node-diagnostics-detail.md § <section>. Catalog of customer-ticket patterns: references/node-issue-catalog.md.


Workflow

  1. Collect cluster name, region, suspect instance ID, exact error string from logs.
  2. Run scripts/triage-cluster.sh (add --node <INSTANCE-ID> to focus one node).
  3. For every [FAIL] / issue entry, Read the referenced section.
  4. Present: what script detected (copy the line verbatim), root cause, exact command(s) with instance/SG IDs filled in, blast radius (e.g. "reboots i-xxx", "wipes volumes on replacement"). For any command that mutates cluster/VPC/SG/EKS state, ask whether the affected resource is IaC-managed and surface the drift warning from the operating-policy note above.
  5. Wait for explicit customer approval. Destructive order: investigate → reboot → replace.
  6. Re-run triage to confirm. Iterate if not cleared.

Step 1: Triage

bash
bash scripts/triage-cluster.sh --cluster <CLUSTER_NAME_OR_ARN> --region <REGION>

# Focus on one node:
bash scripts/triage-cluster.sh --cluster <CLUSTER_NAME_OR_ARN> --region <REGION> --node <INSTANCE_ID>

One pass collects: cluster status + NodeRecovery, events, per-node health (HyperPod + EKS labels, Slurm states), VPC/SG snapshot, CloudWatch availability, SSM readiness, on-node resource checks (disk, memory, /dev/shm, OOM, NVMe, time sync, SSM agent), Slurm node→instance mapping.

Tags: [PASS] passed · [FAIL] issue with a → references/... pointer · [WARN] advisory · [INFO] informational. Priorities: P0 blocks operation · P1 degraded · P2 informational.

Step 2: Match signal → section

Events (list-cluster-events) — provisioning-time:

EventSection
"EFA health checks did not run successfully" (public-doc verbatim signal)A: EFA/SG
Instance bootstrap or network-misconfiguration eventA + B: VPC
Lifecycle-script failure or timeoutD: Lifecycle
Insufficient-capacity or AZ-mismatch failure at creationC: Capacity
Hardware failure / UnschedulablePendingReplacementF: Hardware

EKS labels:

LabelSection
node-health-status: UnschedulablePendingReplacementF
node-health-status: UnschedulablePendingRebootF
deep-health-check-status: FailedG → F

Symptoms:

SymptomSection
Training hangs at NCCL init / AllReduceA → E
Slurm node down / "Node unexpectedly rebooted"H: Slurm
Jobs stuck PENDING / COMPLETINGH
Auto-repair not triggeringF
GPU not visible / XID / ECC errorsG
GPU row-remap pending/failed / silent NaNs / DCGM FailG § G.1.a/b
Disk full / OOM / "Cannot allocate memory"I: Resources
Wrong vCPU count (e.g. 96 instead of 192 on p5.48xlarge)J: Config
Container CrashLoopBackOff / runtime crashM: Container Runtime
aws-node CrashLoopBackOff / gRPC 50051 refusedO: CNI / Pod Networking
Pods stuck Pending with no IP / CNI errorO
DNS resolution / enableDnsSupportB § B.2
Public subnet / IGW misconfiguredB § B.3
Missing VPC endpoints (ECR / STS / FSx)B § B.4
EKS VPC / SG mismatch with HyperPodB § B.5
Kernel panic / watchdog / hung taskN: Kernel
Need shell on a nodeK: SSM
Collect logs for AWS SupportL: Log Collection

A: EFA / Security Group

Per the HyperPod prerequisites doc, the SG must allow all inbound and outbound to itself. scripts/check-efa-sg.sh validates self-ref rules on every cluster SG. On-node EFA check via scripts/check-node-reachability.sh over SSM. Full: § A.

B: VPC / Routing

SG/subnet VPC mismatch, missing S3 Gateway endpoint, EKS auth mode, worker→controller routing, VPC DNS support, private-subnet + NAT / VPC endpoints, EKS↔HyperPod VPC alignment. scripts/check-vpc-config.sh. Full: § B.

C: Capacity / AZ

Insufficient-capacity failure at creation, or no subnets in the AZ where capacity is available. Check AZ offerings via describe-instance-type-offerings, then change subnet AZ or use Flexible Training Plans / ODCR. Full: § C.

D: Lifecycle Scripts

Surfaced in cluster events + CloudWatch under LifecycleConfig/<group>/<instance-id>. Common: S3 connectivity, IAM gaps, CRLF line endings, infinite loops, parameter-name mismatch. Full: § D.

E: Software Versions

Delegate to hyperpod-version-checker to compare NVIDIA driver, CUDA, NCCL, EFA installer, OFI NCCL, PyTorch across nodes. Ensure job env has FI_PROVIDER=efa, FI_EFA_USE_DEVICE_RDMA=1, NCCL_SOCKET_IFNAME=^lo,docker. Full: § E.

F: Hardware / Auto-Repair

Confirm NodeRecovery=Automatic, inspect the EKS health labels + sagemaker.amazonaws.com/fault-details annotation, and read the SagemakerHealthMonitoringAgent/<group>/<instance> CloudWatch stream. HMA runs passive background checks on GPU and Neuron state and reboots the node on count mismatch (per the HMA doc: "if there's a mismatch between the expected number of GPUs … and the count returned by nvidia-smi, then HMA reboots the node"; same for neuron-ls). Manual recovery order: reboot first, replace only if reboot fails; the preferred path is the batch APIs (BatchReboot/BatchReplaceClusterNodes). Full: § F · patterns: node-issue-catalog.md.

G: GPU / Accelerator

NVIDIA (p4d/p5/g5/g6): nvidia-smi + dmesg over SSM for Xid, ECC, thermal throttling. Xid classification per NVIDIA's catalog: 13 Graphics Engine Exception (application-level), 31 GPU memory page fault (application, can be driver/HW), 63 GPU memory remapping event (HW/ECC), 71 CE4 Error (HW copy engine), 74 NVLink Error (HW), 79 GPU has fallen off the bus (PCIe bus), 109 Context Switch Timeout Error (HW). Any uncorrectable ECC → drain and replace. Row-remap state is the authoritative silent-degradation signal (§ G.1.a).

Trainium / Inferentia (trn1/trn2/inf2): Neuron SDK — neuron-ls, neuron-top, neuron-monitor. nvidia-smi does not apply.

GPU / accelerator failures flow into § F for reboot / replace. Full: § G.

H: Slurm Node Management

Node down/unresponsive, unexpected reboots, stuck PENDING/COMPLETING jobs, Slurm-to-instance-ID translation. Primary access is SSM; diagnose slurmd first, fix the root cause, then start/resume the node per § H. Full: § H.

I: Resource Exhaustion

Disk full (HyperPod root volume defaults to 100 GB and is not intended to grow post-creation), OOM, os.fork() memory error, /dev/shm exhaustion, inode exhaustion. Fork-memory fix: export FI_EFA_USE_HUGE_PAGE=0. Redirect bulk data to /opt/sagemaker (secondary EBS) or /opt/dlami/nvme (instance store). Full: § I.

Show full SKILL.md (652 more words)Show less

J: Configuration

p5.48xlarge reports 96 vCPU instead of 192 → set ThreadsPerCore=2 via update-cluster. Full: § J.

K: Node Access via SSM

No direct SSH on HyperPod. Target format sagemaker-cluster:<CLUSTER_ID>_<GROUP>-<INSTANCE_ID>. Failures: plugin missing, wrong prefix, IAM, VPC endpoints. Full: § K.

L: Log Collection

Delegate to hyperpod-issue-report for S3-stored bundles. Key CloudWatch streams: LifecycleConfig/<group>/<instance-id>, SagemakerHealthMonitoringAgent/<group>/<instance-id>. Full: § L.

M: Container Runtime

CrashLoopBackOff, OOMKilled, ImagePullBackOff, RunContainerError on EKS. kubectl describe pod + on-node crictl ps -a, journalctl -u containerd. Full: § M.

N: Kernel & System

Kernel panic, watchdog timeout, soft lockup, unexpected reboots not explained by HyperPod health monitoring. dmesg | grep -iE 'panic|watchdog|hung_task|NMI' + journalctl -b -1. nvrm-related signatures point at NVIDIA driver crashes. Full: § N.

O: CNI / Pod Networking

VPC CNI (aws-node) failures, IPAMD errors, gRPC 127.0.0.1:50051 refused, pods stuck Pending with FailedCreatePodSandBox. Script auto-checks aws-node, kube-proxy, CoreDNS. Full: § O.


Prerequisites

  • aws CLI v2, recent enough to support the HyperPod cluster commands (describe-cluster, list-cluster-nodes, batch-reboot-cluster-nodes, batch-replace-cluster-nodes)
  • python3, bash 4+ (associative arrays are required by the scripts)
  • kubectl authenticated to the EKS cluster (K8s checks skipped if absent)
  • session-manager-plugin for on-node hardware checks
  • unbuffer (from the expect package) — optional; if missing, SSM on-node probes are skipped while the rest of the triage still runs. Install via yum install expect / apt install expect.

Defaults

  • Region — required: pass --region or set $AWS_DEFAULT_REGION.
  • Target scope — all nodes; --node <ID> focuses one.
  • Event window — up to 500 most recent events (5 × 100, paginated).
  • Node list cap — up to 20,000 nodes (200 × 100); warns on cap.
  • SSM probes — 180 s per node with retry-on-throttle.
  • Colors — auto-disabled on non-TTY; --no-color to force off.

Error handling

FailureScriptTell the customer
aws sts get-caller-identity failsExit 1"Fix AWS credentials and rerun."
describe-cluster failsExit 1 after listing region's clusters"Confirm cluster name and region."
sagemaker:* / ec2:* / logs:* AccessDeniedWarn, add Missing IAM permission for <API>, continue"Grant the listed IAM action and rerun."
kubectl absent or unauthenticatedSkip K8s checks"Install/authenticate kubectl (see § K)."
session-manager-plugin absentSkip on-node probes"Install session-manager-plugin (see § K)."
SSM start-session fails or times out (180s)Mark node unreachable with → § K pointer"Rerun with --node <ID> to isolate; verify SSM agent on the node."
Cluster > 20,000 nodesFirst 20,000 paginated; warn"Use --node to target specific nodes."

Exit codes: 0 triage complete · 1 cluster not found or fatal prerequisite missing.

IAM permissions

Read-only diagnostic — covers triage-cluster.sh, check-efa-sg.sh, check-vpc-config.sh, and check-node-reachability.sh:

json
{
  "Action": [
    "sagemaker:DescribeCluster",
    "sagemaker:DescribeClusterNode",
    "sagemaker:ListClusterNodes",
    "sagemaker:ListClusterEvents",
    "sagemaker:ListClusters",
    "eks:DescribeCluster",
    "ec2:DescribeSecurityGroups",
    "ec2:DescribeSubnets",
    "ec2:DescribeVpcs",
    "ec2:DescribeVpcAttribute",
    "ec2:DescribeVpcEndpoints",
    "ec2:DescribeRouteTables",
    "ec2:DescribeNetworkInterfaces",
    "ec2:DescribeInstances",
    "ec2:DescribeInstanceTypeOfferings",
    "ec2:DescribeInstanceTypes",
    "logs:DescribeLogGroups",
    "logs:DescribeLogStreams",
    "logs:FilterLogEvents",
    "ssm:StartSession",
    "ssm:TerminateSession",
    "service-quotas:GetServiceQuota"
  ]
}

sts:GetCallerIdentity is implicit — it requires no IAM action. SSM on HyperPod uses start-session against sagemaker-cluster:<cluster-id>_<group>-<iid> targets — not send-command against bare instance IDs. For remediation commands, grant the matching write permission (e.g. ec2:AuthorizeSecurityGroupIngress / Egress, ec2:RevokeSecurityGroupIngress / Egress, ec2:ModifyVpcAttribute, sagemaker:UpdateCluster, sagemaker:BatchRebootClusterNodes, sagemaker:BatchReplaceClusterNodes). Not needed for the diagnostic itself.

Skill delegation

NeedUse
Cluster creation / deployment failureshyperpod-cluster-debugger (§ A / B / C / H + --validate)
Cluster-wide SSM outagehyperpod-cluster-debugger § F
Single-node SSM failurestay here — § K
Cluster-wide EFA health-check failure at creation timehyperpod-cluster-debugger § A
Single-node EFA failure post-provisioningstay here — § A
NCCL AllReduce / collective-op timeouts (distributed)hyperpod-nccl
Silent GPU NaNs on a specific node (row-remap / DCGM)stay here — § G.1 (even if discovered by NCCL)
Post-deployment cluster-wide managementhyperpod-cluster-debugger
Shell / commands on nodeshyperpod-ssm
CUDA / NCCL / EFA version comparisonhyperpod-version-checker
Diagnostic bundle for AWS Supporthyperpod-issue-report
Training performance / MFU degradationhyperpod-mfu-debugger

Escalate to AWS Support

Escalate when:

  1. SG rules correct and reachability passes but EFA still fails.
  2. VPC correct but K8s bootstrap fails — check VPC flow logs for REJECT.
  3. Hardware failure where replacement keeps failing (bad physical host).
  4. Node replacement fails with an insufficient-capacity signal despite a valid ODCR.
Before opening the case
bash
# 1. Cluster identity + affected node status
aws sagemaker describe-cluster --cluster-name <CLUSTER> --region <REGION>
aws sagemaker list-cluster-nodes --cluster-name <CLUSTER> --region <REGION> \
  --query "ClusterNodeSummaries[?InstanceId=='<INSTANCE_ID>']"

# 2. Triage bundle (scoped to the affected node where possible)
bash scripts/triage-cluster.sh --cluster <CLUSTER> --region <REGION> --node <INSTANCE_ID> > triage.txt

# 3. Per-node log/config bundle to S3 (delegates to hyperpod-issue-report)
#    See skills/hyperpod-issue-report/SKILL.md for the exact invocation.
Include in the case
  • Cluster name + ARN and AWS region
  • Orchestrator (EKS or Slurm)
  • Affected instance IDs / node names / instance-group names
  • Timestamp window (UTC start / end) of the failure
  • Exact error strings observed (copy verbatim from pod logs, CloudWatch, dmesg, events)
  • XID numbers / ECC counts / DCGM output where hardware is implicated
  • triage.txt from step 2 above
  • S3 URI of the hyperpod-issue-report bundle from step 3

Patterns from real customer tickets: node-issue-catalog.md.

© awslabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in plugins/sagemaker-ai/skills/hyperpod-node-debugger of awslabs/agent-plugins.

  • SKILL.md
  • references/node-diagnostics-detail.md
  • references/node-issue-catalog.md
  • scripts/check-efa-sg.sh
  • scripts/check-node-reachability.sh
  • scripts/check-vpc-config.sh
  • scripts/triage-cluster.sh

Open the folder on GitHubat commit da51970

Compare with similar skills

Hyperpod Node Debugger next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hyperpod Node Debugger compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hyperpod Node Debugger this skillawslabs/agent-plugins915—~5.2kAutomated safety check: PassApache-2.0
SkyPilot Multi-Cloud OrchestrationOrchestra-Research/AI-Research-SKILLs13k4 repos~2.4kAutomated safety check: PassMIT
Dstackdstackai/dstack2.3k—~6.2kAutomated safety check: WarnMPL-2.0
Paidf Orchestration SetupNVIDIA/skills3.5k—~3.8kAutomated safety check: WarnApache-2.0
Inspire ML Platform CLIrealZillionX/InspireSkill549—~1.4kAutomated safety check: PassMIT
Areno Debug RuntimeinclusionAI/AReno323—~486Automated safety check: PassApache-2.0

Similar skills

  • SkyPilot Multi-Cloud Orchestration

    Orchestra-Research/AI-Research-SKILLs

    Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost.

    13k GitHub starsUsed in 4 repos~2.4k tokens
    DevOps & CloudAuto-check passed
  • Dstack

    dstackai/dstack

    dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.

    2.3k GitHub stars~6.2k tokensUpdated today
    DevOps & CloudAuto-check: warnings
  • Official

    Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar.

    3.5k GitHub stars~3.8k tokensUpdated today
    DevOps & CloudAuto-check: warnings
  • Inspire ML Platform CLI

    realZillionX/InspireSkill

    Operates the Inspire ML platform through its local `inspire` CLI: picking account, workspace and resources, launching notebooks, jobs and services, then cleaning up.

    549 GitHub stars~1.4k tokensUpdated 10 days ago
    DevOps & CloudAuto-check passed
  • Areno Debug Runtime

    inclusionAI/AReno

    Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno.

    323 GitHub stars~486 tokensUpdated 14 days ago
    AI & LLM EngineeringAuto-check passed
  • TensorRT-LLM Inference

    Orchestra-Research/AI-Research-SKILLs

    Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

    13k GitHub starsUsed in 5 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from awslabs/agent-plugins

All 33 skills in this repo
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    915 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    915 GitHub starsUsed in 2 repos~3.5k tokens
    Auto-check passed
  • Finetuning Technique

    awslabs/agent-plugins

    Official

    Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes.

    915 GitHub starsUsed in 1 repo~604 tokens
    Auto-check passed
  • Hyperpod Issue Report

    awslabs/agent-plugins

    Official

    Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases.

    915 GitHub starsUsed in 1 repo~890 tokens
    Auto-check passed
  • Hyperpod Performance Debugger

    awslabs/agent-plugins

    Official

    Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

    915 GitHub starsUsed in 1 repo~4.1k tokens
    Auto-check passed
  • Hyperpod Ssm

    awslabs/agent-plugins

    Official

    Remote command execution and file transfer on SageMaker HyperPod cluster nodes via AWS Systems Manager (SSM).

    915 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check: notes

Questions about Hyperpod Node Debugger

What does Hyperpod Node Debugger do?

Diagnose and remediate per-node issues on a HyperPod cluster (EKS or Slurm) — a specific node is unhealthy, unresponsive, stuck, or needs replacing. Hyperpod Node Debugger is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose and remediate per-node issues on a HyperPod cluster (EKS or Slurm) — a specific node is unhealthy, unresponsive, stuck, or needs replacing.

When should I use Hyperpod Node Debugger?

Hyperpod Node Debugger fits situations like: tasks that involve GPU and accelerator computing.

How do I install Hyperpod Node Debugger in Claude Code?

Run `npx skills add awslabs/agent-plugins --skill hyperpod-node-debugger -a claude-code`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-node-debugger in awslabs/agent-plugins) into .claude/skills/hyperpod-node-debugger in your project. Claude Code loads it when a task matches its description.

How do I install Hyperpod Node Debugger in Codex?

Run `npx skills add awslabs/agent-plugins --skill hyperpod-node-debugger -a codex`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-node-debugger in awslabs/agent-plugins) into .agents/skills/hyperpod-node-debugger in your project. Codex loads it when a task matches its description.

Can I use Hyperpod Node Debugger in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awslabs/agent-plugins --skill hyperpod-node-debugger -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hyperpod-node-debugger, .gemini/skills/hyperpod-node-debugger, .github/skills/hyperpod-node-debugger and .opencode/skills/hyperpod-node-debugger in your project.

What does Hyperpod Node Debugger need to run?

Going by SKILL.md and its folder, Hyperpod Node Debugger needs a shell for the scripts in its folder and the command-line tools its instructions call (bash, aws, kubectl, yum and apt). Our summary lists: A Bash shell; Docker.

Does Hyperpod Node Debugger access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Hyperpod Node Debugger safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Hyperpod Node Debugger use?

Hyperpod Node Debugger is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hyperpod Node Debugger use?

About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.

What are the alternatives to Hyperpod Node Debugger?

Skills that share tags, products or a category with Hyperpod Node Debugger: SkyPilot Multi-Cloud Orchestration (Orchestra-Research/AI-Research-SKILLs, 13k stars), Dstack (dstackai/dstack, 2.3k stars), Paidf Orchestration Setup (NVIDIA/skills, 3.5k stars) and Inspire ML Platform CLI (realZillionX/InspireSkill, 549 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hyperpod Node Debugger?

awslabs (a GitHub organization, an official publisher) maintains it in awslabs/agent-plugins, which has 915 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 5, 2026.

Source: awslabs/agent-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.