Hyperpod Performance Debugger
awslabs/agent-plugins
Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.
A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.
$ npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aws/tools-for-devops-agent aiml-gpu-training-cluster-investigation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/aiml-gpu-training-cluster-investigation .claude/skills/aiml-gpu-training-cluster-investigation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "aiml-gpu-training-cluster-investigation" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aiml-gpu-training-cluster-investigation into .claude/skills/aiml-gpu-training-cluster-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aiml-gpu-training-cluster-investigation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aws/tools-for-devops-agent/tree/main/skills/aiml-gpu-training-cluster-investigationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aws/tools-for-devops-agent aiml-gpu-training-cluster-investigation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/aiml-gpu-training-cluster-investigation .agents/skills/aiml-gpu-training-cluster-investigation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "aiml-gpu-training-cluster-investigation" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aiml-gpu-training-cluster-investigation into .agents/skills/aiml-gpu-training-cluster-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aiml-gpu-training-cluster-investigation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aws/tools-for-devops-agent aiml-gpu-training-cluster-investigation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/aiml-gpu-training-cluster-investigation .cursor/skills/aiml-gpu-training-cluster-investigation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "aiml-gpu-training-cluster-investigation" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aiml-gpu-training-cluster-investigation into .cursor/skills/aiml-gpu-training-cluster-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aiml-gpu-training-cluster-investigation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aws/tools-for-devops-agent.git --path skills/aiml-gpu-training-cluster-investigation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aws/tools-for-devops-agent aiml-gpu-training-cluster-investigation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/aiml-gpu-training-cluster-investigation .gemini/skills/aiml-gpu-training-cluster-investigation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "aiml-gpu-training-cluster-investigation" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aiml-gpu-training-cluster-investigation into .gemini/skills/aiml-gpu-training-cluster-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aiml-gpu-training-cluster-investigation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aws/tools-for-devops-agent aiml-gpu-training-cluster-investigationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/aiml-gpu-training-cluster-investigation .github/skills/aiml-gpu-training-cluster-investigation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "aiml-gpu-training-cluster-investigation" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aiml-gpu-training-cluster-investigation into .github/skills/aiml-gpu-training-cluster-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aiml-gpu-training-cluster-investigation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aws/tools-for-devops-agent aiml-gpu-training-cluster-investigation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/aiml-gpu-training-cluster-investigation .opencode/skills/aiml-gpu-training-cluster-investigation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "aiml-gpu-training-cluster-investigation" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aiml-gpu-training-cluster-investigation into .opencode/skills/aiml-gpu-training-cluster-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aiml-gpu-training-cluster-investigation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
aiml-gpu-training-cluster-investigationA skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.
Aiml GPU Training Cluster Investigation is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. Use this skill for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances. It adds three things. First, a GPU evidence coverage audit that proves per node whether kernel Xid logs and HyperPod health-agent detections were actually arriving, hour by hour, so "no errors found" is never reported from a silent log. Second, a node verdict (replace, reboot, or leave alone) against an explicit evidence bar, so an application Xid is never headlined…
Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 253 other files, including reference files (for example `.skilleval.yaml`, `CHANGELOG.md` and `README.md`).
It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with Amazon SageMaker and Amazon Web Services. The repository describes itself as: Open-source tools for AWS DevOps Agent - extend DevOps Agent with ready-to-use skills, custom agents, and other tools, for incident response, root cause analysis, and operational…. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ddda70b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.aws.amazon.comdocs.nvidia.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Aiml GPU Training Cluster Investigation loads about 5.4k tokens when it runs, and up to ~26k if it reads all its reference files. Until then it costs about 265 tokens; SKILL.md has 2,590 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from aws/tools-for-devops-agent at commit ddda70b, republished under its Apache-2.0 licence (© aws). 2,590 words, ~5,373 tokens.
.claude/skills/aiml-gpu-training-cluster-investigation/SKILL.md (or your agent's skills folder). This skill also uses 245 other files; get the full folder from GitHub.For GPU clusters on SageMaker HyperPod (Slurm or EKS), AWS ParallelCluster, and self-managed EC2 or EKS. Generic investigation finds the obvious signals; this skill adds what it gets wrong: silent evidence, wrong-headline hardware verdicts, and predictable failures before a long run. Read-only. Never reboot, replace, update, or delete anything, and never read training data, checkpoints, or model weights.
R1. Answer in one pass, and always leave room to answer. In chat, do not stop to ask a
question and do not hand off to a separate investigation before answering. If an input is
missing, use the default (impact window: last 24 hours; run length: evaluate 24, 48, and
72 hours), state the assumption, and mark dependent checks Needs input. For "slow" or
performance questions with no time given, use the last 72 hours. Offer follow-ups only
after the answer.
Budget the evidence gathering so the answer always gets written. An investigation that
runs out of room before it reports is worth nothing to the operator, and it is worse than a
partial answer because it looks like a failure rather than a finding. So: collect the
mandatory evidence for the mode first (R2, R4, R5, and for Mode P the P1 to P6 core), then
write the report. Pick up the optional checks only with what is left. If you notice you are
deep into tool calls and have not yet produced an answer, stop collecting and report what
you have, marking everything unreached as Not checked with the call that would close
it. Never end a turn with evidence gathered and no verdict.
R2. Inventory and capability profile. HyperPod: sagemaker.DescribeCluster and
ListClusterNodes (paginate). Read NodeProvisioningMode from DescribeCluster: if it is
Continuous, also pull sagemaker.ListClusterEvents for the window (see rule R11), which
is the only timeline source that survives broken log delivery. On any other value the call
is unsupported and must be skipped, not retried.
EC2/ParallelCluster/EKS: ec2.DescribeInstances. For every
GPU instance type: ec2.DescribeInstanceTypes (strip HyperPod ml.): GPU count,
EfaSupported, MaximumEfaInterfaces; per EC2 node, attached efa/efa-only interfaces.
Report <attached> of <max>, where attached = interfaces with InterfaceType efa or
efa-only (the primary ENA interface does not count unless it is efa). HyperPod nodes
are not visible to DescribeInstances: say so. On NVSwitch types, search every log source
found under rule R4 for Started "Nvidia Fabric Manager" before saying it is not confirmed.
R3. Node identity survives replacement. A HyperPod reboot keeps the instance ID; a replace
gives the node a new instance ID in the same instance group, so the current ID will
never appear in the replace request. Query CloudTrail by event name, not by instance
ID: cloudtrail.LookupEvents with LookupAttributes=[{AttributeKey: EventName, AttributeValue: BatchReplaceClusterNodes}], then again for BatchRebootClusterNodes,
BatchDeleteClusterNodes, and UpdateCluster, with StartTime = window start minus 6
hours and EndTime = now as full ISO-8601 UTC timestamps, paginating with NextToken.
Keep events whose requestParameters.clusterName is this cluster. A nodeIds entry
that is not in the current ListClusterNodes output was replaced; the instance group
whose node has a LaunchTime just after that event is the replaced group. That operator
or automatic call is the explanation for the node going Pending (Branch E), not hardware.
R4. Find every log source by substring, not prefix. Call logs.DescribeLogGroups with
logGroupNamePattern (case-sensitive substring) = the cluster name, then again for
kernel, messages, syslog, journal, and gpu, paginating with nextToken. Never
search only /aws/parallelcluster or /aws/sagemaker prefixes: customer pipelines use
other names (for example /aws/<pipeline>/<cluster>/kernel). Evaluate every source found.
R5. Prove coverage before any "no errors". For each node and source: find the stream that
carries kernel: lines, then bin that exact stream by hour across the window padded by
one hour. Always name the evidence you used: quote the full log group name and the exact
log stream name for every node in the coverage table, and again in the answer text. A
coverage claim without the group and stream it rests on is not auditable, so the operator
cannot re-run it. Any empty hour means Not observable for that hour. First and last event times
are not proof, and the time of the last kernel: line is not when logging stopped:
a healthy kernel goes quiet. Liveness comes only from the hourly bins of all lines in
that stream. A node is Measured if one source passes. HyperPod: a missing
SagemakerHealthMonitoringAgent/<group>/<instance-id> stream means No HMA detections
when the cluster log group is otherwise live. NCCL transport with no NCCL INFO lines
anywhere is Not observable; never infer it from the instance type.
R5a. Name every resource you looked at, by ID. A finding the operator cannot re-run is
not a finding. Whatever you analysed, put its identifier in the answer: the FSx file system
(fs-...) behind any storage claim, the instance IDs (i-...) behind any node claim, the
cluster name, the capacity reservation (cr-...) behind any capacity claim, and the log
group and stream behind any log claim as R5 already requires. "The file system was
saturated" or "the metrics looked fine" names nothing and cannot be checked. This applies
to the resource you cleared as much as the one you blamed, since ruling something out is
only useful if the reader knows what was ruled out.
R6. Verdict per node, headline to match. REPLACE, REBOOT, LEAVE ALONE, MONITOR, or
NOT OBSERVABLE, against the evidence bar in references/incident-branches.md (Step 4b).
Application-class Xids (for example 13, 31) or HMA reason: XidUserAppError with the node
Running is LEAVE ALONE. Never headline "hardware error" unless the verdict is REPLACE
or REBOOT on hardware grounds.
R7. Label every cause Proven (measured signal on the affected node, before the failure,
nothing competing) or Hypothesis (to validate) with the one confirming measurement. A
spike at the same time is correlation. FSx without a saturated metric is not a proven cause.
Only a Proven cause may be called the root cause, in the headline or in a branch table.
Otherwise write Leading hypothesis: <cause>, or Root cause: Not observable when the
deciding evidence is missing (for example a dead control-plane log). Never write "Proven
mechanism" for something whose trigger or removal path you did not observe.
Utilization metrics from FSx (NetworkThroughputUtilization, DiskIopsUtilization, and
similar) and GPUPowerUtilization are already percent from 0 to 100: a value of 0.9 is
0.9 percent. Quote the raw value with a percent sign.
R8. Recovery questions always state three things: whether automatic node recovery is on
(NodeRecovery), what it does (reboot or replace the node), and that the job resumes
only with checkpoints plus the orchestrator's auto-resume (Slurm on HyperPod:
srun --auto-resume=1).
R9. Capacity Blocks begin terminating instances 30 minutes before the end time (60 for
UltraServers); blocks end at 11:30 UTC and termination starts at 11:00 UTC on the last day.
For a planned run, write out: usable until = end time minus the lead time; run end = start
plus run length; hours covered = usable until minus start. Give every value as a full UTC
date and time, and check the latest safe start is not already in the past.
R10. Rule out the frequent non-GPU causes in references/cluster-edge-cases.md before
blaming hardware: subnet IP or network interface exhaustion, ParallelCluster bootstrap
failures and protected mode, EFA nodes in a public subnet, a Capacity Block not yet
active, and the FSx maintenance window. HyperPod does not export system metrics to
CloudWatch, so HyperPod GPU activity is Not observable there.
R11. When the logs are dead, ask the control plane. On a HyperPod cluster with
NodeProvisioningMode = Continuous, sagemaker.ListClusterEvents gives you a node and
cluster timeline that owes nothing to a log agent, so it keeps answering when a stream has
gone silent or a node has disappeared. Filter the window with EventTimeAfter and
EventTimeBefore, narrow with NodeId or InstanceGroupName, sort with
SortBy=EventTime, and page through NextToken. Where a Description is not
self-explanatory, DescribeClusterEvent has the detail. Note that the response has no
severity or level field at all, so any grouping you apply is your own and should be
described that way. If NodeProvisioningMode is anything other than Continuous the call
is not supported; write ListClusterEvents not supported in the coverage table and carry
on. What you must not do is report a dead log as "no events" without either trying this
source or saying it was unavailable.
| The user asks | Mode | Steps to run |
|---|---|---|
| Something failed, hung, slowed, or lost nodes | I: Incident | Steps 1 to 7 |
| "Were there GPU errors?", "Can I trust the logs?" | C: Coverage audit | Steps 1 to 3, then 6 and 7 |
| "Is the cluster ready for a long run?", Capacity Block ending | P: Pre-flight | Steps 1 to 3, then 5P, 6 and 7 |
Work through these in order and tick each one as it completes. Skip only the steps the
mode table excludes. Every step below has a matching ## Step N section with its detail.
Account, region, cluster name or instance IDs, workload, impact window (default last 24 hours, stated). Classify the symptom to pick a starting branch (Step 5), but collect evidence for all.
Load references/inventory-and-timeline.md for the inventory API calls and the eight timeline sources, and references/cluster-edge-cases.md for the frequent non-GPU causes to rule out under rule R10:
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/inventory-and-timeline.md")
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/cluster-edge-cases.md")Build one ordered timeline for the window plus 30 minutes each side: node state, HMA detections, Xids, AWS Health, EC2 status and scheduled events, Capacity Block and training plan end times, and CloudTrail cluster changes (rule R3).
Load references/coverage-audit.md for the log-source discovery and hourly coverage queries, and references/nccl-nvlink-efa.md for NCCL transport, NVLink and NVSwitch, and EFA signals:
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/coverage-audit.md")
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/nccl-nvlink-efa.md")Produce the coverage table and the node capability and fabric table for every affected node. Every row names the full log group name and the exact log stream name that row's verdict rests on, so the operator can re-run the same query. Where no stream carries kernel lines, say which groups you searched and that none did.
Load references/xid-triage.md for the Xid catalog and per-code verdicts, and references/incident-branches.md for the node verdict evidence bar and branches A to F:
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/xid-triage.md")
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/incident-branches.md")Load references/signals-and-thresholds.md for metric names, dimensions, and thresholds:
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/signals-and-thresholds.md")Pull FSx (correct dimensions per metric), GPU activity (AWS/EC2 GPUPowerUtilization, unit
Percent, or CWAgent), and EFA counters, then evaluate branches A (hardware), B (capacity
lifecycle), C (storage), D (NCCL, NVLink/NVSwitch, EFA), E (cluster change), and F (application,
only after A to E are ruled out), as defined in incident-branches.md. Recommend operator
actions only.
Load references/preflight.md for pre-flight checks P1 to P16:
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/preflight.md")Score checks P1 to P16. Lead with FAIL, then RISK. In Mode P this step replaces Steps 4 and 5.
Work the core first, then extend. All sixteen checks together cost more tool calls than a single answer usually has room for, and a readiness question with no verdict is a failed answer however much evidence sits behind it (see R1). So run them in two passes.
The core, which decides whether the run can start at all:
| Check | Question it settles |
|---|---|
| P1 | Does the Capacity Block or training plan outlast the run? |
| P2 | Is there an extension, if it does not? |
| P3 | Is there a spare node to replace a failure? |
| P4 | Is NodeRecovery on? |
| P5 | Are deep health checks enabled? |
| P6 | Is GPU error logging arriving, so a failure during the run is visible? |
Write the readiness verdict as soon as those six are scored. P7 to P16 then refine it, and
each one you reach can only add a RISK, never change a FAIL already found in the core.
Anything you do not reach is reported Not checked with the call that would settle it, which
is an honest answer; silence is not. If the core itself is incomplete, say which part and
give the verdict you can support.
Load references/report-format.md for the report template and its rules:
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/report-format.md")Before showing the answer to the user, re-read your own draft and verify each of these. Fix the draft where a check fails; do not present an output that fails one.
Not observable, not healthy.SagemakerHealthMonitoringAgent/<instance-group>/<instance-id>. This is the easiest
check to skip in a short answer and the one that most often makes a finding
unreproducible.REPLACE or REBOOT on hardware grounds (rule R6).Proven or Hypothesis (to validate) label, and anything
labelled Proven has a measured signal on the affected node before the failure
(rule R7). Nothing unproven is called the root cause.Not observable with what to collect, never as
zero or as healthy.fs-... behind a storage
claim, the i-... behind a node claim, the cr-... behind a capacity claim, the
cluster name, the log group and stream. This holds for resources you cleared, not just
the one you blamed.Not checked (rule R1).State the outcome of this self-check in one line, naming anything you could not verify.
Proven or Hypothesis (to validate).© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 245 other files (references) in skills/aiml-gpu-training-cluster-investigation of aws/tools-for-devops-agent.
Open the folder on GitHubat commit ddda70b
Aiml GPU Training Cluster Investigation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Aiml GPU Training Cluster Investigation this skillaws/tools-for-devops-agent | 103 | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | |
| Hyperpod Performance Debuggerawslabs/agent-plugins | 916 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | |
| Hyperpod Version Checkerawslabs/agent-plugins | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | |
| Hyperpod Ncclawslabs/agent-plugins | 916 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Python Environment Setup for SageMakerhuggingface/skills | 11k | 2 repos | ~1.7k | Automated safety check: Pass | Apache-2.0 |
awslabs/agent-plugins
Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
awslabs/agent-plugins
Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.
awslabs/agent-plugins
Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters.
aws/tools-for-devops-agent
ALWAYS use this skill in the beginning of any incident investigation, root cause analysis, or operational troubleshooting.
aws/tools-for-devops-agent
AWS Database Migration Service (DMS) operational review and troubleshooting skill.
aws/tools-for-devops-agent
Performs a comprehensive Amazon ECS operations review across the 6 review pillars (Resiliency & HA, Observability, Security, Operations, Performance, Additional Analysis) using read-only AWS APIs…
aws/tools-for-devops-agent
Comprehensive Amazon RDS and Aurora operational review aligned with the AWS Well-Architected Framework and RDS/Aurora best practices.
aws/tools-for-devops-agent
Amazon SageMaker AI Operational Review. An agent skill from aws/tools-for-devops-agent.
aws/tools-for-devops-agent
Use this skill during any incident investigation, capacity planning, or operational troubleshooting when the issue may be caused by hitting AWS service limits.
Works with
Categories
A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances. Aiml GPU Training Cluster Investigation is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. Use this skill for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.
Aiml GPU Training Cluster Investigation fits situations like: inference clusters on SageMaker HyperPod (Slurm; parallelCluster; self-managed EC2/EKS GPU instances.
Run `npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a claude-code`. Or copy the skill folder (skills/aiml-gpu-training-cluster-investigation in aws/tools-for-devops-agent) into .claude/skills/aiml-gpu-training-cluster-investigation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a codex`. Or copy the skill folder (skills/aiml-gpu-training-cluster-investigation in aws/tools-for-devops-agent) into .agents/skills/aiml-gpu-training-cluster-investigation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aiml-gpu-training-cluster-investigation, .gemini/skills/aiml-gpu-training-cluster-investigation, .github/skills/aiml-gpu-training-cluster-investigation and .opencode/skills/aiml-gpu-training-cluster-investigation in your project.
SKILL.md names no scripts, command-line tools or credentials: Aiml GPU Training Cluster Investigation is instructions for the agent only.
SKILL.md names 2 domains. As links in the text: docs.aws.amazon.com and docs.nvidia.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Aiml GPU Training Cluster Investigation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.4k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Aiml GPU Training Cluster Investigation: Hyperpod Performance Debugger (awslabs/agent-plugins, 916 stars), Hyperpod Version Checker (awslabs/agent-plugins, 916 stars), Hyperpod Nccl (awslabs/agent-plugins, 916 stars) and SageMaker Serving Image Selection (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aws (a GitHub organization, an official publisher) maintains it in aws/tools-for-devops-agent, which has 103 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 9, 2026.
Source: aws/tools-for-devops-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.