Official agent skill

Hyperpod Performance Debugger

by awslabs in awslabs/agent-plugins

Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Hyperpod Performance Debugger

skills CLI
$ npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install awslabs/agent-plugins hyperpod-performance-debugger --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-performance-debugger .claude/skills/hyperpod-performance-debugger && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hyperpod-performance-debugger
GitHub stars
912
Token cost
~4.1k tokens
SKILL.md length
1,449 words
Files
3 (incl. scripts, references)
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

  • Works in 2 steps: Run the snapshot → Match signal → section
  • Uneven NCCL across nodes
  • SKILL.md covers Scope and delegation, Operating policy, Workflow and Step 1: Run the snapshot, plus 5 more sections
  • Runs Shell scripts from its folder; calls aws and bash

What it does

Hyperpod Performance Debugger is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput. Read-only. Surfaces host-side signals (Xid, ECC, NVLink, EFA reachability, FSx saturation) and routes to the appropriate sibling skill (hyperpod-node-debugger, hyperpod-nccl, hyperpod-version-checker, hyperpod-issue-report) for any remediation. Triggers on uneven NCCL across nodes, straggler node, FSx slow, checkpoint slow, dataloader slow, filesystem bottleneck, FSx throughput…

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/perf-details.md` and `scripts/perf-snapshot.sh`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with Amazon SageMaker and Amazon Web Services. The repository describes itself as: Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS. The licence is Apache-2.0.

When your agent uses it

  • Uneven NCCL across nodes
  • Checkpoint slow
  • Dataloader slow
  • Filesystem bottleneck

Example prompts

  • “/hyperpod-performance-debugger”

Requirements

  • A Bash shell

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Run the snapshot
  2. Match signal → section

What it can do on your machine

Read from SKILL.md and the folder at commit da51970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • aws
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.aws.amazon.com
    • github.com
    • awslabs.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hyperpod Performance Debugger loads about 4.1k tokens when it runs, and up to ~7.1k if it reads all its reference files. Until then it costs about 146 tokens; SKILL.md has 1,449 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from awslabs/agent-plugins at commit da51970, republished under its Apache-2.0 licence (© awslabs). 1,449 words, ~4,115 tokens.

Download SKILL.mdSave it as .claude/skills/hyperpod-performance-debugger/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
hyperpod-performance-debugger
description
Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput. Read-only. Surfaces host-side signals (Xid, ECC, NVLink, EFA reachability, FSx saturation) and routes to the appropriate sibling skill (hyperpod-node-debugger, hyperpod-nccl, hyperpod-version-checker, hyperpod-issue-report) for any remediation. Triggers on uneven NCCL across nodes, straggler node, FSx slow, checkpoint slow, dataloader slow, filesystem bottleneck, FSx throughput, cross-AZ latency, topology mismatch.
metadata.version
0.0.1

HyperPod Performance Debugger

  1. Uneven NCCL performance across nodes — workload faster on some node sets than others, pairwise bandwidth variance, suspected straggler.
  2. Poor filesystem performance — training stalled on data loading, checkpoint save/load dominating step time, FSx throughput saturated.

Scope and delegation

Route findings outside the two in-scope scenarios to the owner skill below.

Concern observedRoute to
GPU hardware fault, ECC, NVLink, Xid, DCGM diagnostics, drain/replacehyperpod-node-debugger (§ F Hardware/Auto-Repair, § G GPU)
Cannot allocate memory at os.fork(), root volume exhaustedhyperpod-node-debugger (§ I Resource Exhaustion)
NCCL timeouts, hangs, AllReduce stalls, EFA TCP fallback, RDMA memlockhyperpod-nccl
EFA / NCCL / CUDA / NVIDIA driver version drift across nodeshyperpod-version-checker
EFA self-referencing security-group rule missing — single nodehyperpod-node-debugger § A (EFA / Security Group)
EFA self-referencing security-group rule missing — cluster-widehyperpod-cluster-debugger § A (EFA Health Checks)
Slurm node state changes (drain / resume / reboot)hyperpod-slurm-debugger
Diagnostic bundle for AWS Supporthyperpod-issue-report
Shell access on a nodehyperpod-ssm

Operating policy

  • Read-only. Print commands the customer runs; do not execute commands that modify state.
  • Container vs host version comparisons go through hyperpod-version-checker.
  • Xid lines, ECC counts, NVLink lane state, and thermal readings get surfaced; the catalog and verdict live in hyperpod-node-debugger § G.

Workflow

  1. Confirm the symptom is uneven NCCL or poor filesystem performance. If neither, route to the matching sibling skill above.
  2. Run scripts/perf-snapshot.sh (read-only) to gather host-side signals for the suspect node and FSx filesystems mounted on it.
  3. For each [CONCERN] line in the script output, open the matching section below and read the supporting reference.
  4. After the per-incident diagnosis, recommend the HyperPod platform health features in § Continuous health coverage so the customer gets ongoing protection.

Step 1: Run the snapshot

bash
bash scripts/perf-snapshot.sh --cluster <CLUSTER_NAME_OR_ARN> --region <REGION>

# Scope to one suspect node:
bash scripts/perf-snapshot.sh --cluster <C> --region <R> --node <INSTANCE_ID>

The script samples one node by default. It collects host-side data via hyperpod-ssm: nvidia-smi output (temperature, SM clocks, PCIe link width, ECC, NVLink, topo -m), recent dmesg Xid lines, EFA port state and fi_info provider visibility, EFA installer + kernel module versions, CPU governor, NVL72 Fabric Manager state, FSx CloudWatch utilization, df -h / lfs df -h per mount, host iowait, /dev/shm size, and root-volume usage. All read-only.

Tags: [OK] healthy · [CONCERN] signal worth investigating (carries a → pointer to the owner skill) · [INFO] informational.

Host vs container scope. The script runs on the host via SSM and reports host-scope values. Many setups ship the EFA / libfabric / OFI-NCCL / CUDA stack inside the training container by design — a host value of unknown is not by itself a defect. What matters for performance is the stack the workload actually uses. Verify versions inside the container (and across nodes) via hyperpod-version-checker before drawing conclusions.

Step 2: Match signal → section

ObservationSection
Pairwise NCCL bandwidth varies across node pairs / suspected stragglerA: Uneven NCCL Performance
Nodes spread across AZs / network-node-layer labels / UltraServer boundariesA
EFA port not ACTIVE on a node, missing OFI plugin, or FI provider not visibleA + route to hyperpod-node-debugger § A; hyperpod-version-checker for cross-node version compare
iostat shows high iowait, FSx CloudWatch utilization sustained near 100%B: Poor Filesystem Performance
DataLoader stalls, checkpoint dominates step timeB
Xid line in dmesg, uncorrectable ECC, inactive NVLink lane, GPU ≥ 88°CRoute to hyperpod-node-debugger § G
Container vs host version drift suspectedRoute to hyperpod-version-checker
Cannot allocate memory at os.fork(), root volume full, OOM eventsRoute to hyperpod-node-debugger § I
NCCL timeout, hang, TCP fallback (NET/OFI Using TCP), RDMA memlockRoute to hyperpod-nccl

A: Uneven NCCL Performance

The customer reports identical training jobs running with different step times on different node sets, pairwise bandwidth variance, or some allocations consistently slower than others despite identical code.

Per the official troubleshooting guide, the common contributing factors are network topology differences between nodes (cross-AZ, cross-rack, cross-UltraServer), degraded EFA performance on some nodes, mixed instance types or generations within an instance group, and CPU frequency scaling differences.

Diagnostic pass (read-only)

The host-side data points — GPU thermal/ECC/PCIe/clocks, Xid, NVLink lanes, EFA port state and provider visibility, CPU governor, EFA/OFI/driver versions, nvidia-smi topo -m — are all collected by scripts/perf-snapshot.sh (Step 1 above). The script tags [CONCERN] with thresholds and emits routing pointers; rerun it per suspect node via --node <INSTANCE_ID>.

For driver / CUDA / NCCL / EFA / OFI version drift across nodes, run hyperpod-version-checker skill.

Pairwise NCCL bandwidth test

Run the standard nccl-tests recipes from awslabs/awsome-distributed-training. For an N-node cluster, run all-reduce across every pair and record busbw for each pair. Pairs more than ~5% below the run mean (the threshold the AWS validation script flags) are problematic candidates.

Expected busbw per SKU is published in the AI-on-HyperPod NCCL test guide. Benchmark the specific instance type before relying on a number.

Pairwise scripts, HyperPod topology surfaces (HyperPod API, EKS labels, Slurm topology.conf), and GB200 NVL72 specifics are in references/perf-details.md § Uneven NCCL.

Topology verification

HyperPod exposes topology through three operator-visible surfaces:

  • HyperPod API: aws sagemaker describe-cluster-node returns NodeDetails.Placement.AvailabilityZone / AvailabilityZoneId and NodeDetails.UltraServerInfo.Id (UltraServer SKUs only).
  • EKS labels: topology.kubernetes.io/zone, topology.k8s.aws/network-node-layer-{1,2,3} (highest-numbered = closest to instance), topology.k8s.aws/ultraserver-id.
  • Slurm: HyperPod auto-generates topology.conf. Inspect via scontrol show topology.

Tightly coupled work shares the same AZ, the same highest-numbered network-node-layer label (EKS) or the same Slurm topology block, and — for NVL72 jobs — the same UltraServerInfo.Id / topology.k8s.aws/ultraserver-id. If the cluster is spread across AZs or layers, topology must be re-established at provisioning time. Route provisioning changes to hyperpod-cluster-debugger § B (Capacity & AZ).


Show full SKILL.md (589 more words)Show less

B: Poor Filesystem Performance

The customer reports training bottlenecked on data loading, checkpoint save/load dominating step time, executables/scripts loading slowly, or iowait high.

Per the official troubleshooting guide, the resolution path follows this order:

  1. Check CloudWatch metrics on the filesystem.
  2. Check the provisioned performance configuration against workload requirements.
  3. Investigate which operations are causing the I/O — workload demand vs inefficient pattern.
  4. Consider upgrading provisioned performance.
  5. Choose the filesystem type that matches the I/O pattern.

This skill covers steps 1–3. Steps 4–5 are customer decisions; surface the data and let the customer pick.

Diagnostic pass (read-only)

scripts/perf-snapshot.sh (Step 1 above) covers the on-node side of this pass: it discovers FSx mounts, calls aws cloudwatch get-metric-statistics on DataReadBytes and (for OpenZFS) FileServerDiskIopsUtilization, prints df -h for /fsx /opt/dlami/nvme /opt/sagemaker, runs lfs df -h per Lustre mount, and reports iostat iowait. It tags [CONCERN] when OpenZFS IOPS utilization sustains ≥ 80% or iowait > 20%.

For longer windows or additional metrics (DataWriteBytes, Lustre DiskIopsUtilization, OpenZFS FileServerDiskThroughputUtilization), drive the query directly:

bash
aws cloudwatch get-metric-statistics --region <REGION> \
  --namespace AWS/FSx --metric-name DataReadBytes \
  --dimensions Name=FileSystemId,Value=<FSID> \
  --start-time "$(date -u -d '3 hours ago' +%Y-%m-%dT%H:%M:%S)" \
  --end-time   "$(date -u +%Y-%m-%dT%H:%M:%S)" \
  --period 60 --statistics Sum Maximum

The full per-filesystem-type metric catalog is in references/perf-details.md § Filesystem.

Branches

Provisioned capacity is saturated. CloudWatch utilization sustained near 100% across the workload window. Customer decision: scale up the filesystem.

  • FSx for Lustre throughput scales with StorageCapacity × PerUnitStorageThroughput; capacity changes are non-disruptive.
  • FSx for OpenZFS — increase provisioned IOPS or throughput.

I/O pattern is inefficient. CloudWatch shows headroom but the workload is still I/O-bound. Customer decision: change the application.

  • DataLoader: raise num_workers, set pin_memory=True, persistent_workers=True.
  • Checkpointing: use async + sharded (torch.distributed.checkpoint.async_save plus FSDP SHARDED_STATE_DICT). FULL_STATE_DICT serializes through rank 0 and is a frequent root cause.
  • Small-file workloads: Lustre is optimized for large sequential I/O. For millions of small files, use WebDataset / tar shards, FSx for OpenZFS, or NVMe scratch.

Filesystem-selection guidance and the async-checkpoint pattern are in references/perf-details.md § Filesystem.


Continuous health coverage

Once the immediate incident is diagnosed, recommend HyperPod's built-in health features so problems are caught before the next training run rather than after another customer-reported regression.

  • Enable NodeRecovery=Automatic on the cluster. The Health Monitoring Agent (HMA) continuously monitors GPU- and Trainium-based instances and marks instances unhealthy on detected failure. With auto-recovery enabled, HyperPod reboots or replaces the node — no operator intervention.

  • Enable OnStartDeepHealthChecks on every GPU instance group with both check categories:

    • InstanceStress — stress-ng on CPU/memory/disk, GPU and PCI device count verification, DCGM level-4 diagnostics (memory test included), and EFA loopback bandwidth/latency.
    • InstanceConnectivity — multi-node NCCL all-reduce.

    Every newly provisioned or auto-replaced node passes the same hardware bar before accepting jobs.

  • Run on-demand deep health checks when this skill or any sibling surfaces a hardware concern but the cluster is mid-workload. aws sagemaker start-cluster-health-check runs the same checks against a specific instance group; nodes are placed in a Slurm maintenance reservation and the check is queued until any running job completes (not preempted). Console: HyperPod → Clusters → Instances → Run deep health checks.

    Not supported when NodeProvisioningMode=Continuous; one on-demand request per cluster at a time. Requires the latest AMI — run UpdateClusterSoftware first.

Logs land in CloudWatch at /aws/sagemaker/Clusters/<cluster_name>/<cluster_id> under DeepHealthCheckResults/<log_stream_id>, and on each node at /var/log/aws/clusters/sagemaker-deep-health-check.log.

References

  • references/perf-details.md — pairwise NCCL test recipes, HyperPod topology check, GB200 NVL72 placement; CloudWatch metric catalog per filesystem type, async-checkpoint pattern, filesystem selection guide.

External:

© awslabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in plugins/sagemaker-ai/skills/hyperpod-performance-debugger of awslabs/agent-plugins.

  • SKILL.md
  • references/perf-details.md
  • scripts/perf-snapshot.sh

Open the folder on GitHubat commit da51970

Compare with similar skills

Hyperpod Performance Debugger next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hyperpod Performance Debugger compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hyperpod Performance Debugger this skillawslabs/agent-plugins912—~4.1kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Python Environment Setup for SageMakerhuggingface/skills11k2 repos~1.7kAutomated safety check: PassApache-2.0
SageMaker Deployment Plannerhuggingface/skills11k1 repos~2.1kAutomated safety check: PassApache-2.0
Hf Cloud Serving Image Selectionwaybarrios/opencode-power-pack533—~4.3kAutomated safety check: PassApache-2.0
AWS AI MLaws/agent-toolkit-for-aws2.8k—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.

    11k GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hf Cloud Serving Image Selection

    waybarrios/opencode-power-pack

    Select and verify the current region-specific serving container URI for a SageMaker model deployment.

    533 GitHub stars~4.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • AWS AI ML

    aws/agent-toolkit-for-aws

    Official

    Selects, deploys, and customizes AI models on Amazon SageMaker.

    2.8k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.

    11k GitHub starsUsed in 1 repo~1.8k tokens
    DevOps & CloudAuto-check passed

More from awslabs/agent-plugins

All 33 skills in this repo
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    912 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    912 GitHub starsUsed in 2 repos~3.5k tokens
    Auto-check passed
  • Finetuning Technique

    awslabs/agent-plugins

    Official

    Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes.

    912 GitHub starsUsed in 1 repo~604 tokens
    Auto-check passed
  • AWS Lambda Managed Instances

    awslabs/agent-plugins

    Official

    Evaluate, configure, and migrate workloads to AWS Lambda Managed Instances (LMI).

    912 GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed
  • Hyperpod Issue Report

    awslabs/agent-plugins

    Official

    Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases.

    912 GitHub stars~890 tokensUpdated 2 days ago
    Auto-check passed
  • Hyperpod Slurm Debugger

    awslabs/agent-plugins

    Official

    Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters.

    912 GitHub stars~3.3k tokensUpdated 2 days ago
    Auto-check passed

Questions about Hyperpod Performance Debugger

What does Hyperpod Performance Debugger do?

Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput. Hyperpod Performance Debugger is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

When should I use Hyperpod Performance Debugger?

Hyperpod Performance Debugger fits situations like: uneven NCCL across nodes; checkpoint slow; dataloader slow; filesystem bottleneck.

How do I install Hyperpod Performance Debugger in Claude Code?

Run `npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a claude-code`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-performance-debugger in awslabs/agent-plugins) into .claude/skills/hyperpod-performance-debugger in your project. Claude Code loads it when a task matches its description.

How do I install Hyperpod Performance Debugger in Codex?

Run `npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a codex`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-performance-debugger in awslabs/agent-plugins) into .agents/skills/hyperpod-performance-debugger in your project. Codex loads it when a task matches its description.

Can I use Hyperpod Performance Debugger in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hyperpod-performance-debugger, .gemini/skills/hyperpod-performance-debugger, .github/skills/hyperpod-performance-debugger and .opencode/skills/hyperpod-performance-debugger in your project.

What does Hyperpod Performance Debugger need to run?

Going by SKILL.md and its folder, Hyperpod Performance Debugger needs a shell for the scripts in its folder and the command-line tools its instructions call (aws and bash). Our summary lists: A Bash shell.

Does Hyperpod Performance Debugger access the network?

SKILL.md names 3 domains. As links in the text: docs.aws.amazon.com, github.com and awslabs.github.io. This is read from the text; nothing was executed.

Is Hyperpod Performance Debugger safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Hyperpod Performance Debugger use?

Hyperpod Performance Debugger is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hyperpod Performance Debugger use?

About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Hyperpod Performance Debugger?

Skills that share tags, products or a category with Hyperpod Performance Debugger: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Python Environment Setup for SageMaker (huggingface/skills, 11k stars), SageMaker Deployment Planner (huggingface/skills, 11k stars) and Hf Cloud Serving Image Selection (waybarrios/opencode-power-pack, 533 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hyperpod Performance Debugger?

awslabs (a GitHub organization, an official publisher) maintains it in awslabs/agent-plugins, which has 912 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 5, 2026.

Source: awslabs/agent-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.