Official agent skill

Hyperpod Slurm Debugger

by awslabs in awslabs/agent-plugins

Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Hyperpod Slurm Debugger

skills CLI
$ npx skills add awslabs/agent-plugins --skill hyperpod-slurm-debugger -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install awslabs/agent-plugins hyperpod-slurm-debugger --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-slurm-debugger .claude/skills/hyperpod-slurm-debugger && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hyperpod-slurm-debugger
GitHub stars
915
Token cost
~3.3k tokens
SKILL.md length
1,021 words
Files
3 (incl. scripts, references)
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters.

  • Works in 4 steps: Collect inputs → Confirm orchestrator → Run the diagnostic script → …
  • Reports a Slurm node stuck in down/drain
  • SKILL.md covers When to invoke, When NOT to invoke, Constraints and Prerequisites, plus 10 more sections
  • Runs Shell scripts from its folder; calls aws, bash and node

What it does

Hyperpod Slurm Debugger is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters. Scope mirrors the HyperPod troubleshooting guide. Invoke when the user reports a Slurm node stuck in down/drain, "Node unexpectedly rebooted" after auto-repair, slurmd not running, jobs stuck PENDING with REASON=Resources while sinfo shows idle nodes, jobs stuck COMPLETING after node replacement, GRES/GPU counts wrong, scontrol ping failing, slurmctld unresponsive, an Action:Reboot/Replace request that…

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/slurm-details.md` and `scripts/slurm-diagnose.sh`).

It sits in AI & LLM Engineering. It works with Amazon SageMaker and Amazon Web Services. The repository describes itself as: Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS. The licence is Apache-2.0.

When your agent uses it

  • Reports a Slurm node stuck in down/drain
  • Node unexpectedly rebooted after auto-repair
  • Slurmd not running
  • Jobs stuck PENDING with REASON=Resources while sinfo shows idle nodes

Example prompts

  • “Node unexpectedly rebooted”
  • “drain before reboot”
  • “diagnose a Slurm node”
  • “/hyperpod-slurm-debugger”

Requirements

  • A Bash shell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Collect inputs
  2. Confirm orchestrator
  3. Run the diagnostic script
  4. Map findings → docs

What it can do on your machine

Read from SKILL.md and the folder at commit da51970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • aws
    • bash
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • slurm.schedmd.com
    • docs.aws.amazon.com
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hyperpod Slurm Debugger loads about 3.3k tokens when it runs, and up to ~6.7k if it reads all its reference files. Until then it costs about 177 tokens; SKILL.md has 1,021 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~177
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from awslabs/agent-plugins at commit da51970, republished under its Apache-2.0 licence (© awslabs). 1,021 words, ~3,262 tokens.

Download SKILL.mdSave it as .claude/skills/hyperpod-slurm-debugger/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
hyperpod-slurm-debugger
description
Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters. Scope mirrors the HyperPod troubleshooting guide. Invoke when the user reports a Slurm node stuck in down/drain, "Node unexpectedly rebooted" after auto-repair, slurmd not running, jobs stuck PENDING with REASON=Resources while sinfo shows idle nodes, jobs stuck COMPLETING after node replacement, GRES/GPU counts wrong, scontrol ping failing, slurmctld unresponsive, an Action:Reboot/Replace request that did not trigger HyperPod auto-recovery, or auto-resume not restarting a job. Also triggers on "drain before reboot", "diagnose a Slurm node", "investigate stuck jobs."
metadata.version
0.0.1

HyperPod Slurm Debugger

Diagnostic-only. Identify and classify Slurm scheduler and node-daemon issues on HyperPod Slurm clusters. Do not run, recommend, or print any state-mutating command. For remediation, link to the official AWS or Slurm documentation.

When to invoke

Invoke when the user reports any of the symptoms in the decision table.

When NOT to invoke

  • Cluster has Orchestrator.Eks — invoke hyperpod-node-debugger or hyperpod-nccl.
  • Single-node hardware fault with healthy Slurm scheduler — invoke hyperpod-node-debugger.
  • NCCL training-hang investigation — invoke hyperpod-nccl.
  • Node unreachable via SSM — invoke hyperpod-ssm.

Constraints

  • Read-only. Do not run, recommend, or print state-mutating commands.
  • For any remediation, link to AWS or Slurm docs. The user authorizes and executes.
  • IaC-managed cluster (Terraform / CloudFormation / CDK): warn that direct mutation drifts the live state from the IaC plan.

Canonical recovery URLs: references/slurm-details.md → Authoritative recovery documentation.

Prerequisites

  • AWS CLI v2, authenticated for the target account and region with permissions:
    • sagemaker:DescribeCluster, sagemaker:ListClusterNodes
    • ssm:StartSession on the HyperPod-created SSM document
  • Session Manager plugin installed locally.
  • jq ≥ 1.6.
  • unbuffer (from the expect package). Required — without it aws ssm start-session returns empty stdout intermittently with Cannot perform start session: EOF and every check silently misreports. Install: expect package on Amazon Linux / RHEL / Debian / Ubuntu / macOS. Script exits at prerequisite check if missing.

Procedure

Step 1 — Collect inputs

Ask the user for:

  1. HyperPod cluster name (not Slurm partition name).
  2. AWS region.
  3. Optional: a specific Slurm node name.
Step 2 — Confirm orchestrator
bash
aws sagemaker describe-cluster --cluster-name <NAME/ARN> --region <REGION> \
  --query 'Orchestrator' --output json

If Orchestrator.Eks is present, stop. Route per When NOT to invoke.

Step 3 — Run the diagnostic script
bash
bash scripts/slurm-diagnose.sh --cluster <NAME> --region <REGION>
# Scope to a node:
bash scripts/slurm-diagnose.sh --cluster <NAME> --region <REGION> --node <SLURM_NODE>

Relay the script output to the user verbatim.

Step 4 — Map findings → docs

For each finding, look up the section in the decision table and link the user to the corresponding AWS / Slurm doc. Do not type out remediation commands.

Decision table

Symptom (sinfo -o "%N %T %30E" or script finding)Section
Node state = down or down*, reason other than belowA: Node Down
Node state = down*, Reason = Node unexpectedly rebootedB: Unexpected Reboot
Jobs PENDING with REASON=Resources while nodes are idleC: Controller State
Jobs stuck COMPLETING after node replacementC: Controller State
scontrol ping returns DOWN for the controllerC: Controller State
GRES (GPU) counts incorrect or not releasedC: Controller State
state=fail issued but no recovery occurredD: Action Reason Mismatch
Accounting errors or RPC errors mentioning dbdC: Controller State (slurmdbd)
slurm.conf edited; new partitions or nodes not visibleC: Controller State (config)
Job exited on a hardware failure but did not restartE: Auto-resume

Defaults

BehaviorDefaultOverride
Moderead-only — always; no remediation flag existsn/a
Region$AWS_DEFAULT_REGION, falling back to us-east-1--region <R>
Scopeall nodes in down / drain / fail / "unexpectedly rebooted"--node <SLURM_NODE_NAME>
Outputcolorized terminal--no-color
SSM target formatsagemaker-cluster:<clusterId>_<instanceGroupName>-<instanceId> (derived)n/a
Controller discovery--controller-group (if set) → SlurmConfig.NodeType=Controller → provisioning_parameters.json--controller-group <N>

Error handling

FailureSkill behaviorRequired user action
describe-cluster failsPrint AWS error; exit 1Fix credentials/region; verify cluster name
Cluster has Orchestrator.EksExit 1 with pointer to EKS-side skillsUse hyperpod-node-debugger or hyperpod-nccl
session-manager-plugin missing / SSM unreachablesinfo returns empty; exit 1Install plugin; verify node InService
Disk ≥ 95 % full on a down nodeReport finding disk-full-<node>Refer to AWS troubleshooting docs
Missing jq or awsExit 1 at prerequisite checkInstall per Prerequisites

A: Node Down

Node is down because slurmd stopped responding. Causes: slurmd crash, disk full, OOM, network partition, hardware fault.

Script checks: systemctl is-active slurmd, srun -w <NODE> hostname (RPC layer), disk, memory.

Link: https://github.com/aws/sagemaker-hyperpod-cluster-setup/blob/troubleshooting-doc-20250917/troubleshoot/index.md

If node returns to down after a manual resume → escalate to hyperpod-node-debugger.

Context: references/slurm-details.md § A.


B: Unexpected Reboot

Node is down* with Reason "Node unexpectedly rebooted" because slurmd re-registered after an out-of-band reboot. Upstream Slurm behavior, not HyperPod. Node is typically healthy.

Links:

If node reboots again within minutes → escalate to hyperpod-node-debugger.

Context: references/slurm-details.md § B.


Show full SKILL.md (394 more words)Show less

C: Controller State

slurmctld in-memory state can desync from the on-disk state. A controller restart reloads from StateSaveLocation and clears bad caches. User decides and executes.

Restart may help:

SymptomWhy
PENDING with REASON=Resources, idle nodesRe-evaluates the queue
Jobs stuck COMPLETING after node replacementController held a reference to the old node
GRES (GPU, EFA) not released after a job endsResource accounting de-synced
Nodes stuck Unknown after reboot, slurmd is upRe-registration was not processed
scontrol ping times outController event loop is hung
Lost connection to slurmdbd / RPC errorsDBD connection wedged

Do NOT restart when:

  • HyperPod replacement (Action:Replace) in progress on any node — concurrent changes fail the replacement.
  • Only one compute node is bad — restart slurmd on that node.
  • sinfo and squeue are responsive — problem is elsewhere.
  • journalctl -u slurmctld not reviewed yet — panic / OOM will reproduce.
  • slurm.conf was just edited — try scontrol reconfigure first.
Folded triggers

Restart procedure / what's preserved:

Context: references/slurm-details.md § C.


D: Action Reason Mismatch

scontrol update state=fail reason=... was issued with a reason that does not match Action:Reboot or Action:Replace exactly. HyperPod silently ignores anything else. Script detects near-misses on nodes in fail state.

Required strings (case-sensitive, no whitespace, no punctuation):

  • Action:Reboot
  • Action:Replace

Link: https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-resiliency-slurm-replace-faulty-instance.html

Context: references/slurm-details.md § Action reason-string validation.


E: Auto-resume

--auto-resume=1 is an srun step option. It re-runs the step after HMA (the Health Monitoring Agent) flags a node and Automatic node recovery replaces it.

Why it didn't restart the job:

  • Flag on sbatch not srun — per-step; sbatch directives are silently ignored.
  • HMA did not flag the node — failure was application/transient, not hardware. Step exits as a normal Slurm failure.
  • Cluster NodeRecovery is None — faulty nodes are labeled but not replaced.
  • No checkpointing — step restarts from process zero each iteration.
  • AMI predates HMA support (released 2025-09-11) — needs AMI / cluster-software update.

Link: https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-resiliency-slurm-auto-resume.html

Context: references/slurm-details.md § HyperPod auto-resume.


Escalation

ConditionNext skill
Node returns to down shortly after a manual resumehyperpod-node-debugger (hardware)
slurmd logs contain CUDA / NVIDIA / XID errorshyperpod-node-debugger § G
Disk full or /dev/shm exhaustedhyperpod-node-debugger § I
Node unreachable via SSMhyperpod-ssm
Controller restart does not clear COMPLETING after 2 attemptshyperpod-issue-report + AWS Support

© awslabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in plugins/sagemaker-ai/skills/hyperpod-slurm-debugger of awslabs/agent-plugins.

  • SKILL.md
  • references/slurm-details.md
  • scripts/slurm-diagnose.sh

Open the folder on GitHubat commit da51970

Compare with similar skills

Hyperpod Slurm Debugger next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hyperpod Slurm Debugger compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hyperpod Slurm Debugger this skillawslabs/agent-plugins915—~3.3kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Python Environment Setup for SageMakerhuggingface/skills11k2 repos~1.7kAutomated safety check: PassApache-2.0
Aiml GPU Training Cluster Investigationaws/tools-for-devops-agent100—~5.4kAutomated safety check: PassApache-2.0
SageMaker Deployment Plannerhuggingface/skills11k1 repos~2.1kAutomated safety check: PassApache-2.0
Hf Cloud Serving Image Selectionwaybarrios/opencode-power-pack533—~4.3kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Aiml GPU Training Cluster Investigation

    aws/tools-for-devops-agent

    Official

    A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.

    100 GitHub stars~5.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.

    11k GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hf Cloud Serving Image Selection

    waybarrios/opencode-power-pack

    Select and verify the current region-specific serving container URI for a SageMaker model deployment.

    533 GitHub stars~4.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • AWS AI ML

    aws/agent-toolkit-for-aws

    Official

    Selects, deploys, and customizes AI models on Amazon SageMaker.

    2.8k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from awslabs/agent-plugins

All 33 skills in this repo
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    915 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    915 GitHub starsUsed in 2 repos~3.5k tokens
    Auto-check passed
  • Finetuning Technique

    awslabs/agent-plugins

    Official

    Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes.

    915 GitHub starsUsed in 1 repo~604 tokens
    Auto-check passed
  • Hyperpod Issue Report

    awslabs/agent-plugins

    Official

    Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases.

    915 GitHub starsUsed in 1 repo~890 tokens
    Auto-check passed
  • Hyperpod Performance Debugger

    awslabs/agent-plugins

    Official

    Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

    915 GitHub starsUsed in 1 repo~4.1k tokens
    Auto-check passed
  • Hyperpod Ssm

    awslabs/agent-plugins

    Official

    Remote command execution and file transfer on SageMaker HyperPod cluster nodes via AWS Systems Manager (SSM).

    915 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check: notes

Questions about Hyperpod Slurm Debugger

What does Hyperpod Slurm Debugger do?

Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters. Hyperpod Slurm Debugger is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters.

When should I use Hyperpod Slurm Debugger?

Hyperpod Slurm Debugger fits situations like: reports a Slurm node stuck in down/drain; Node unexpectedly rebooted after auto-repair; slurmd not running; jobs stuck PENDING with REASON=Resources while sinfo shows idle nodes.

How do I install Hyperpod Slurm Debugger in Claude Code?

Run `npx skills add awslabs/agent-plugins --skill hyperpod-slurm-debugger -a claude-code`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-slurm-debugger in awslabs/agent-plugins) into .claude/skills/hyperpod-slurm-debugger in your project. Claude Code loads it when a task matches its description.

How do I install Hyperpod Slurm Debugger in Codex?

Run `npx skills add awslabs/agent-plugins --skill hyperpod-slurm-debugger -a codex`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-slurm-debugger in awslabs/agent-plugins) into .agents/skills/hyperpod-slurm-debugger in your project. Codex loads it when a task matches its description.

Can I use Hyperpod Slurm Debugger in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awslabs/agent-plugins --skill hyperpod-slurm-debugger -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hyperpod-slurm-debugger, .gemini/skills/hyperpod-slurm-debugger, .github/skills/hyperpod-slurm-debugger and .opencode/skills/hyperpod-slurm-debugger in your project.

What does Hyperpod Slurm Debugger need to run?

Going by SKILL.md and its folder, Hyperpod Slurm Debugger needs a shell for the scripts in its folder and the command-line tools its instructions call (aws, bash and node). Our summary lists: A Bash shell.

Does Hyperpod Slurm Debugger access the network?

SKILL.md names 3 domains. As links in the text: slurm.schedmd.com, docs.aws.amazon.com and github.com. This is read from the text; nothing was executed.

Is Hyperpod Slurm Debugger safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Hyperpod Slurm Debugger use?

Hyperpod Slurm Debugger is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hyperpod Slurm Debugger use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.

What are the alternatives to Hyperpod Slurm Debugger?

Skills that share tags, products or a category with Hyperpod Slurm Debugger: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Python Environment Setup for SageMaker (huggingface/skills, 11k stars), Aiml GPU Training Cluster Investigation (aws/tools-for-devops-agent, 100 stars) and SageMaker Deployment Planner (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hyperpod Slurm Debugger?

awslabs (a GitHub organization, an official publisher) maintains it in awslabs/agent-plugins, which has 915 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 5, 2026.

Source: awslabs/agent-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.