Official agent skill

Aiml GPU Training Cluster Investigation

by aws in aws/tools-for-devops-agent

A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Aiml GPU Training Cluster Investigation

skills CLI
$ npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aws/tools-for-devops-agent aiml-gpu-training-cluster-investigation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/aiml-gpu-training-cluster-investigation .claude/skills/aiml-gpu-training-cluster-investigation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aiml-gpu-training-cluster-investigation
GitHub stars
103
Token cost
~5.4k tokens
SKILL.md length
2,590 words
Files
246 (incl. references)
Skills in repo
31
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.

  • Works in 7 steps: Scope → Inventory and timeline → Coverage audit → …
  • Inference clusters on SageMaker HyperPod (Slurm
  • SKILL.md covers Critical rules R1 to R11…, Pick the mode, Workflow checklist and Step 1: Scope, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Aiml GPU Training Cluster Investigation is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. Use this skill for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances. It adds three things. First, a GPU evidence coverage audit that proves per node whether kernel Xid logs and HyperPod health-agent detections were actually arriving, hour by hour, so "no errors found" is never reported from a silent log. Second, a node verdict (replace, reboot, or leave alone) against an explicit evidence bar, so an application Xid is never headlined…

Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 253 other files, including reference files (for example `.skilleval.yaml`, `CHANGELOG.md` and `README.md`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with Amazon SageMaker and Amazon Web Services. The repository describes itself as: Open-source tools for AWS DevOps Agent - extend DevOps Agent with ready-to-use skills, custom agents, and other tools, for incident response, root cause analysis, and operational…. The licence is Apache-2.0.

When your agent uses it

  • Inference clusters on SageMaker HyperPod (Slurm
  • ParallelCluster
  • Self-managed EC2/EKS GPU instances

Example prompts

  • “no errors found”
  • “is my cluster ready for a multi-day run”
  • “/aiml-gpu-training-cluster-investigation”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Scope
  2. Inventory and timeline
  3. Coverage audit
  4. Classify faults and give node verdicts
  5. Metrics and root-cause branch
  6. Report
  7. Self-check before presenting

What it can do on your machine

Read from SKILL.md and the folder at commit ddda70b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.aws.amazon.com
    • docs.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Aiml GPU Training Cluster Investigation loads about 5.4k tokens when it runs, and up to ~26k if it reads all its reference files. Until then it costs about 265 tokens; SKILL.md has 2,590 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~265
When it runs · the whole SKILL.md, loaded when a task matches
~5.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~26k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aws/tools-for-devops-agent at commit ddda70b, republished under its Apache-2.0 licence (© aws). 2,590 words, ~5,373 tokens.

Download SKILL.mdSave it as .claude/skills/aiml-gpu-training-cluster-investigation/SKILL.md (or your agent's skills folder). This skill also uses 245 other files; get the full folder from GitHub.
name
aiml-gpu-training-cluster-investigation
description
Use this skill for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances. It adds three things. First, a GPU evidence coverage audit that proves per node whether kernel Xid logs and HyperPod health-agent detections were actually arriving, hour by hour, so "no errors found" is never reported from a silent log. Second, a node verdict (replace, reboot, or leave alone) against an explicit evidence bar, so an application Xid is never headlined as hardware. Third, a pre-flight readiness check before a long run, covering Capacity Block or training plan end time versus run length, spare capacity to replace a failed node, NodeRecovery, deep health checks, EFA, log coverage, idle reserved GPUs. Activate on Xid or ECC errors, slow training or FSx for Lustre slowness on a GPU cluster, NCCL hangs or TCP fallback, NVLink or Fabric Manager errors, nodes in Failure or Pending, nodes terminating at once, or "is my cluster ready for a multi-day run".
metadata.author
nzuresh
metadata.version
1.0.4
metadata.aws-devops-agent-skills.agent-t
Incident RCA, Chat tasks
metadata.aws-devops-agent-skills.aws-ser
Amazon SageMaker HyperPod, AWS ParallelCluster, Amazon EC2, Amazon FSx for Lustre, Elastic Fabric Adapter, Amazon EKS, AWS Health
metadata.aws-devops-agent-skills.technic
Machine Learning, GenAI, High Performance Computing

GPU Cluster Evidence, Readiness, and Fault Verdicts

For GPU clusters on SageMaker HyperPod (Slurm or EKS), AWS ParallelCluster, and self-managed EC2 or EKS. Generic investigation finds the obvious signals; this skill adds what it gets wrong: silent evidence, wrong-headline hardware verdicts, and predictable failures before a long run. Read-only. Never reboot, replace, update, or delete anything, and never read training data, checkpoints, or model weights.

Critical rules R1 to R11 (apply in every mode, in this order)

R1. Answer in one pass, and always leave room to answer. In chat, do not stop to ask a question and do not hand off to a separate investigation before answering. If an input is missing, use the default (impact window: last 24 hours; run length: evaluate 24, 48, and 72 hours), state the assumption, and mark dependent checks Needs input. For "slow" or performance questions with no time given, use the last 72 hours. Offer follow-ups only after the answer. Budget the evidence gathering so the answer always gets written. An investigation that runs out of room before it reports is worth nothing to the operator, and it is worse than a partial answer because it looks like a failure rather than a finding. So: collect the mandatory evidence for the mode first (R2, R4, R5, and for Mode P the P1 to P6 core), then write the report. Pick up the optional checks only with what is left. If you notice you are deep into tool calls and have not yet produced an answer, stop collecting and report what you have, marking everything unreached as Not checked with the call that would close it. Never end a turn with evidence gathered and no verdict. R2. Inventory and capability profile. HyperPod: sagemaker.DescribeCluster and ListClusterNodes (paginate). Read NodeProvisioningMode from DescribeCluster: if it is Continuous, also pull sagemaker.ListClusterEvents for the window (see rule R11), which is the only timeline source that survives broken log delivery. On any other value the call is unsupported and must be skipped, not retried. EC2/ParallelCluster/EKS: ec2.DescribeInstances. For every GPU instance type: ec2.DescribeInstanceTypes (strip HyperPod ml.): GPU count, EfaSupported, MaximumEfaInterfaces; per EC2 node, attached efa/efa-only interfaces. Report <attached> of <max>, where attached = interfaces with InterfaceType efa or efa-only (the primary ENA interface does not count unless it is efa). HyperPod nodes are not visible to DescribeInstances: say so. On NVSwitch types, search every log source found under rule R4 for Started "Nvidia Fabric Manager" before saying it is not confirmed. R3. Node identity survives replacement. A HyperPod reboot keeps the instance ID; a replace gives the node a new instance ID in the same instance group, so the current ID will never appear in the replace request. Query CloudTrail by event name, not by instance ID: cloudtrail.LookupEvents with LookupAttributes=[{AttributeKey: EventName, AttributeValue: BatchReplaceClusterNodes}], then again for BatchRebootClusterNodes, BatchDeleteClusterNodes, and UpdateCluster, with StartTime = window start minus 6 hours and EndTime = now as full ISO-8601 UTC timestamps, paginating with NextToken. Keep events whose requestParameters.clusterName is this cluster. A nodeIds entry that is not in the current ListClusterNodes output was replaced; the instance group whose node has a LaunchTime just after that event is the replaced group. That operator or automatic call is the explanation for the node going Pending (Branch E), not hardware. R4. Find every log source by substring, not prefix. Call logs.DescribeLogGroups with logGroupNamePattern (case-sensitive substring) = the cluster name, then again for kernel, messages, syslog, journal, and gpu, paginating with nextToken. Never search only /aws/parallelcluster or /aws/sagemaker prefixes: customer pipelines use other names (for example /aws/<pipeline>/<cluster>/kernel). Evaluate every source found. R5. Prove coverage before any "no errors". For each node and source: find the stream that carries kernel: lines, then bin that exact stream by hour across the window padded by one hour. Always name the evidence you used: quote the full log group name and the exact log stream name for every node in the coverage table, and again in the answer text. A coverage claim without the group and stream it rests on is not auditable, so the operator cannot re-run it. Any empty hour means Not observable for that hour. First and last event times are not proof, and the time of the last kernel: line is not when logging stopped: a healthy kernel goes quiet. Liveness comes only from the hourly bins of all lines in that stream. A node is Measured if one source passes. HyperPod: a missing SagemakerHealthMonitoringAgent/<group>/<instance-id> stream means No HMA detections when the cluster log group is otherwise live. NCCL transport with no NCCL INFO lines anywhere is Not observable; never infer it from the instance type. R5a. Name every resource you looked at, by ID. A finding the operator cannot re-run is not a finding. Whatever you analysed, put its identifier in the answer: the FSx file system (fs-...) behind any storage claim, the instance IDs (i-...) behind any node claim, the cluster name, the capacity reservation (cr-...) behind any capacity claim, and the log group and stream behind any log claim as R5 already requires. "The file system was saturated" or "the metrics looked fine" names nothing and cannot be checked. This applies to the resource you cleared as much as the one you blamed, since ruling something out is only useful if the reader knows what was ruled out. R6. Verdict per node, headline to match. REPLACE, REBOOT, LEAVE ALONE, MONITOR, or NOT OBSERVABLE, against the evidence bar in references/incident-branches.md (Step 4b). Application-class Xids (for example 13, 31) or HMA reason: XidUserAppError with the node Running is LEAVE ALONE. Never headline "hardware error" unless the verdict is REPLACE or REBOOT on hardware grounds. R7. Label every cause Proven (measured signal on the affected node, before the failure, nothing competing) or Hypothesis (to validate) with the one confirming measurement. A spike at the same time is correlation. FSx without a saturated metric is not a proven cause. Only a Proven cause may be called the root cause, in the headline or in a branch table. Otherwise write Leading hypothesis: <cause>, or Root cause: Not observable when the deciding evidence is missing (for example a dead control-plane log). Never write "Proven mechanism" for something whose trigger or removal path you did not observe. Utilization metrics from FSx (NetworkThroughputUtilization, DiskIopsUtilization, and similar) and GPUPowerUtilization are already percent from 0 to 100: a value of 0.9 is 0.9 percent. Quote the raw value with a percent sign. R8. Recovery questions always state three things: whether automatic node recovery is on (NodeRecovery), what it does (reboot or replace the node), and that the job resumes only with checkpoints plus the orchestrator's auto-resume (Slurm on HyperPod: srun --auto-resume=1). R9. Capacity Blocks begin terminating instances 30 minutes before the end time (60 for UltraServers); blocks end at 11:30 UTC and termination starts at 11:00 UTC on the last day. For a planned run, write out: usable until = end time minus the lead time; run end = start plus run length; hours covered = usable until minus start. Give every value as a full UTC date and time, and check the latest safe start is not already in the past. R10. Rule out the frequent non-GPU causes in references/cluster-edge-cases.md before blaming hardware: subnet IP or network interface exhaustion, ParallelCluster bootstrap failures and protected mode, EFA nodes in a public subnet, a Capacity Block not yet active, and the FSx maintenance window. HyperPod does not export system metrics to CloudWatch, so HyperPod GPU activity is Not observable there. R11. When the logs are dead, ask the control plane. On a HyperPod cluster with NodeProvisioningMode = Continuous, sagemaker.ListClusterEvents gives you a node and cluster timeline that owes nothing to a log agent, so it keeps answering when a stream has gone silent or a node has disappeared. Filter the window with EventTimeAfter and EventTimeBefore, narrow with NodeId or InstanceGroupName, sort with SortBy=EventTime, and page through NextToken. Where a Description is not self-explanatory, DescribeClusterEvent has the detail. Note that the response has no severity or level field at all, so any grouping you apply is your own and should be described that way. If NodeProvisioningMode is anything other than Continuous the call is not supported; write ListClusterEvents not supported in the coverage table and carry on. What you must not do is report a dead log as "no events" without either trying this source or saying it was unavailable.

Pick the mode

The user asksModeSteps to run
Something failed, hung, slowed, or lost nodesI: IncidentSteps 1 to 7
"Were there GPU errors?", "Can I trust the logs?"C: Coverage auditSteps 1 to 3, then 6 and 7
"Is the cluster ready for a long run?", Capacity Block endingP: Pre-flightSteps 1 to 3, then 5P, 6 and 7

Workflow checklist

Work through these in order and tick each one as it completes. Skip only the steps the mode table excludes. Every step below has a matching ## Step N section with its detail.

  • Step 1: Scope the request: account, region, cluster or instance IDs, impact window
  • Step 2: Build the inventory, capability profile, and one ordered timeline
  • Step 3: Prove GPU log coverage per node before looking for errors
  • Step 4: Classify each fault and give every node a verdict
  • Step 5: Pull metrics and settle the root-cause branch
  • Step 5P: Score pre-flight checks P1 to P16 (Mode P only, replaces Steps 4 and 5)
  • Step 6: Write the report in the required format
  • Step 7: Self-check the finished output, then present it
Show full SKILL.md (1,016 more words)Show less

Step 1: Scope

Account, region, cluster name or instance IDs, workload, impact window (default last 24 hours, stated). Classify the symptom to pick a starting branch (Step 5), but collect evidence for all.

Step 2: Inventory and timeline

Load references/inventory-and-timeline.md for the inventory API calls and the eight timeline sources, and references/cluster-edge-cases.md for the frequent non-GPU causes to rule out under rule R10:

read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/inventory-and-timeline.md")
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/cluster-edge-cases.md")

Build one ordered timeline for the window plus 30 minutes each side: node state, HMA detections, Xids, AWS Health, EC2 status and scheduled events, Capacity Block and training plan end times, and CloudTrail cluster changes (rule R3).

Step 3: Coverage audit

Load references/coverage-audit.md for the log-source discovery and hourly coverage queries, and references/nccl-nvlink-efa.md for NCCL transport, NVLink and NVSwitch, and EFA signals:

read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/coverage-audit.md")
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/nccl-nvlink-efa.md")

Produce the coverage table and the node capability and fabric table for every affected node. Every row names the full log group name and the exact log stream name that row's verdict rests on, so the operator can re-run the same query. Where no stream carries kernel lines, say which groups you searched and that none did.

Step 4: Classify faults and give node verdicts

Load references/xid-triage.md for the Xid catalog and per-code verdicts, and references/incident-branches.md for the node verdict evidence bar and branches A to F:

read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/xid-triage.md")
read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/incident-branches.md")

Step 5: Metrics and root-cause branch

Load references/signals-and-thresholds.md for metric names, dimensions, and thresholds:

read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/signals-and-thresholds.md")

Pull FSx (correct dimensions per metric), GPU activity (AWS/EC2 GPUPowerUtilization, unit Percent, or CWAgent), and EFA counters, then evaluate branches A (hardware), B (capacity lifecycle), C (storage), D (NCCL, NVLink/NVSwitch, EFA), E (cluster change), and F (application, only after A to E are ruled out), as defined in incident-branches.md. Recommend operator actions only.

Step 5P: Pre-flight readiness (Mode P)

Load references/preflight.md for pre-flight checks P1 to P16:

read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/preflight.md")

Score checks P1 to P16. Lead with FAIL, then RISK. In Mode P this step replaces Steps 4 and 5.

Work the core first, then extend. All sixteen checks together cost more tool calls than a single answer usually has room for, and a readiness question with no verdict is a failed answer however much evidence sits behind it (see R1). So run them in two passes.

The core, which decides whether the run can start at all:

CheckQuestion it settles
P1Does the Capacity Block or training plan outlast the run?
P2Is there an extension, if it does not?
P3Is there a spare node to replace a failure?
P4Is NodeRecovery on?
P5Are deep health checks enabled?
P6Is GPU error logging arriving, so a failure during the run is visible?

Write the readiness verdict as soon as those six are scored. P7 to P16 then refine it, and each one you reach can only add a RISK, never change a FAIL already found in the core. Anything you do not reach is reported Not checked with the call that would settle it, which is an honest answer; silence is not. If the core itself is incomplete, say which part and give the verdict you can support.

Step 6: Report

Load references/report-format.md for the report template and its rules:

read_skill_resource(skill_id="aiml-gpu-training-cluster-investigation", path="references/report-format.md")

Step 7: Self-check before presenting

Before showing the answer to the user, re-read your own draft and verify each of these. Fix the draft where a check fails; do not present an output that fails one.

  • Every "no errors found" statement is backed by a node whose coverage you proved in Step 3. If coverage was not proven, the wording is Not observable, not healthy.
  • Every coverage row names its full log group and exact log stream (rule R5). A coverage claim with no named source is not auditable and must be fixed before presenting.
  • Stream names appear as the service writes them, not paraphrased. Search your own draft for phrases like "the HMA log stream" or "the health agent log" and replace each with the real name, for example SagemakerHealthMonitoringAgent/<instance-group>/<instance-id>. This is the easiest check to skip in a short answer and the one that most often makes a finding unreproducible.
  • Every node verdict still meets the evidence bar that justifies it, re-read from references/incident-branches.md Step 4b.
  • The headline matches the verdicts. It does not say "hardware error" unless a verdict is REPLACE or REBOOT on hardware grounds (rule R6).
  • Every cause carries a Proven or Hypothesis (to validate) label, and anything labelled Proven has a measured signal on the affected node before the failure (rule R7). Nothing unproven is called the root cause.
  • Every percentage came straight from the metric without rescaling (rule R7).
  • Every absent signal is reported as Not observable with what to collect, never as zero or as healthy.
  • Each recommendation names an operator action, and no mutating API call was made.
  • Every number in the answer can be traced to a call you actually made this run.
  • Every resource you analysed appears by ID (rule R5a): the fs-... behind a storage claim, the i-... behind a node claim, the cr-... behind a capacity claim, the cluster name, the log group and stream. This holds for resources you cleared, not just the one you blamed.
  • There is an actual answer. A verdict or root cause is written down, not just evidence. If you ran out of room before finishing, the draft still leads with the verdict you can support and marks the rest Not checked (rule R1).

State the outcome of this self-check in one line, naming anything you could not verify.

Success criteria

  • Coverage table and node capability table for every affected node; no "no errors" without proven coverage.
  • One verdict per node with a GPU signal; headline consistent with the verdicts.
  • Every cause labelled Proven or Hypothesis (to validate).
  • Replaced nodes matched to the operator or automatic action that replaced them.
  • Mode P: P1 to P16 scored.
  • No mutating API call was made.

References

© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 245 other files (references) in skills/aiml-gpu-training-cluster-investigation of aws/tools-for-devops-agent.

  • SKILL.md
  • .skilleval.yaml
  • CHANGELOG.md
  • README.md
  • evals/best-practices/v10/benchmark.json
  • evals/best-practices/v10/iteration-1/best-practices-tests-results.json
  • evals/best-practices/v10/iteration-2/best-practices-tests-results.json
  • evals/best-practices/v10/iteration-3/best-practices-tests-results.json
  • evals/evals.json
  • evals/functional/v5/_metadata.json
  • evals/functional/v5/benchmark.json
  • evals/functional/v5/evals.json
  • evals/functional/v5/iteration-1
  • … and 233 more

Open the folder on GitHubat commit ddda70b

Compare with similar skills

Aiml GPU Training Cluster Investigation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Aiml GPU Training Cluster Investigation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Aiml GPU Training Cluster Investigation this skillaws/tools-for-devops-agent103—~5.4kAutomated safety check: PassApache-2.0
Hyperpod Performance Debuggerawslabs/agent-plugins916—~4.1kAutomated safety check: PassApache-2.0
Hyperpod Version Checkerawslabs/agent-plugins916—~910Automated safety check: PassApache-2.0
Hyperpod Ncclawslabs/agent-plugins916—~3.4kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Python Environment Setup for SageMakerhuggingface/skills11k2 repos~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Hyperpod Performance Debugger

    awslabs/agent-plugins

    Official

    Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

    916 GitHub stars~4.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    916 GitHub stars~910 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Nccl

    awslabs/agent-plugins

    Official

    Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…

    916 GitHub stars~3.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Slurm Debugger

    awslabs/agent-plugins

    Official

    Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters.

    916 GitHub stars~3.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from aws/tools-for-devops-agent

All 31 skills in this repo
  • AWS Health Events

    aws/tools-for-devops-agent

    Official

    ALWAYS use this skill in the beginning of any incident investigation, root cause analysis, or operational troubleshooting.

    103 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Database Migration Service Expertise

    aws/tools-for-devops-agent

    Official

    AWS Database Migration Service (DMS) operational review and troubleshooting skill.

    103 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ecs Operation Review

    aws/tools-for-devops-agent

    Official

    Performs a comprehensive Amazon ECS operations review across the 6 review pillars (Resiliency & HA, Observability, Security, Operations, Performance, Additional Analysis) using read-only AWS APIs…

    103 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Rds Operation Review

    aws/tools-for-devops-agent

    Official

    Comprehensive Amazon RDS and Aurora operational review aligned with the AWS Well-Architected Framework and RDS/Aurora best practices.

    103 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Sagemaker AI Ops Review

    aws/tools-for-devops-agent

    Official

    Amazon SageMaker AI Operational Review. An agent skill from aws/tools-for-devops-agent.

    103 GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Service Quota Check

    aws/tools-for-devops-agent

    Official

    Use this skill during any incident investigation, capacity planning, or operational troubleshooting when the issue may be caused by hitting AWS service limits.

    103 GitHub stars~3.4k tokensUpdated today
    Auto-check passed

Questions about Aiml GPU Training Cluster Investigation

What does Aiml GPU Training Cluster Investigation do?

A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances. Aiml GPU Training Cluster Investigation is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. Use this skill for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.

When should I use Aiml GPU Training Cluster Investigation?

Aiml GPU Training Cluster Investigation fits situations like: inference clusters on SageMaker HyperPod (Slurm; parallelCluster; self-managed EC2/EKS GPU instances.

How do I install Aiml GPU Training Cluster Investigation in Claude Code?

Run `npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a claude-code`. Or copy the skill folder (skills/aiml-gpu-training-cluster-investigation in aws/tools-for-devops-agent) into .claude/skills/aiml-gpu-training-cluster-investigation in your project. Claude Code loads it when a task matches its description.

How do I install Aiml GPU Training Cluster Investigation in Codex?

Run `npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a codex`. Or copy the skill folder (skills/aiml-gpu-training-cluster-investigation in aws/tools-for-devops-agent) into .agents/skills/aiml-gpu-training-cluster-investigation in your project. Codex loads it when a task matches its description.

Can I use Aiml GPU Training Cluster Investigation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/tools-for-devops-agent --skill aiml-gpu-training-cluster-investigation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aiml-gpu-training-cluster-investigation, .gemini/skills/aiml-gpu-training-cluster-investigation, .github/skills/aiml-gpu-training-cluster-investigation and .opencode/skills/aiml-gpu-training-cluster-investigation in your project.

What does Aiml GPU Training Cluster Investigation need to run?

SKILL.md names no scripts, command-line tools or credentials: Aiml GPU Training Cluster Investigation is instructions for the agent only.

Does Aiml GPU Training Cluster Investigation access the network?

SKILL.md names 2 domains. As links in the text: docs.aws.amazon.com and docs.nvidia.com. This is read from the text; nothing was executed.

Is Aiml GPU Training Cluster Investigation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Aiml GPU Training Cluster Investigation use?

Aiml GPU Training Cluster Investigation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Aiml GPU Training Cluster Investigation use?

About 5.4k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.

What are the alternatives to Aiml GPU Training Cluster Investigation?

Skills that share tags, products or a category with Aiml GPU Training Cluster Investigation: Hyperpod Performance Debugger (awslabs/agent-plugins, 916 stars), Hyperpod Version Checker (awslabs/agent-plugins, 916 stars), Hyperpod Nccl (awslabs/agent-plugins, 916 stars) and SageMaker Serving Image Selection (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Aiml GPU Training Cluster Investigation?

aws (a GitHub organization, an official publisher) maintains it in aws/tools-for-devops-agent, which has 103 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 9, 2026.

Source: aws/tools-for-devops-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.