SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.
$ npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install awslabs/agent-plugins hyperpod-performance-debugger --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-performance-debugger .claude/skills/hyperpod-performance-debugger && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "hyperpod-performance-debugger" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-performance-debugger into .claude/skills/hyperpod-performance-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-performance-debugger", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-performance-debuggerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install awslabs/agent-plugins hyperpod-performance-debugger --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-performance-debugger .agents/skills/hyperpod-performance-debugger && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "hyperpod-performance-debugger" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-performance-debugger into .agents/skills/hyperpod-performance-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-performance-debugger", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install awslabs/agent-plugins hyperpod-performance-debugger --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-performance-debugger .cursor/skills/hyperpod-performance-debugger && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "hyperpod-performance-debugger" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-performance-debugger into .cursor/skills/hyperpod-performance-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-performance-debugger", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/awslabs/agent-plugins.git --path plugins/sagemaker-ai/skills/hyperpod-performance-debugger--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install awslabs/agent-plugins hyperpod-performance-debugger --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-performance-debugger .gemini/skills/hyperpod-performance-debugger && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "hyperpod-performance-debugger" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-performance-debugger into .gemini/skills/hyperpod-performance-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-performance-debugger", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install awslabs/agent-plugins hyperpod-performance-debuggerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-performance-debugger .github/skills/hyperpod-performance-debugger && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "hyperpod-performance-debugger" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-performance-debugger into .github/skills/hyperpod-performance-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-performance-debugger", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install awslabs/agent-plugins hyperpod-performance-debugger --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-performance-debugger .opencode/skills/hyperpod-performance-debugger && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "hyperpod-performance-debugger" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-performance-debugger into .opencode/skills/hyperpod-performance-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-performance-debugger", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
hyperpod-performance-debuggerDiagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.
Hyperpod Performance Debugger is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput. Read-only. Surfaces host-side signals (Xid, ECC, NVLink, EFA reachability, FSx saturation) and routes to the appropriate sibling skill (hyperpod-node-debugger, hyperpod-nccl, hyperpod-version-checker, hyperpod-issue-report) for any remediation. Triggers on uneven NCCL across nodes, straggler node, FSx slow, checkpoint slow, dataloader slow, filesystem bottleneck, FSx throughput…
Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/perf-details.md` and `scripts/perf-snapshot.sh`).
It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with Amazon SageMaker and Amazon Web Services. The repository describes itself as: Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS. The licence is Apache-2.0.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit da51970. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
awsbashFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.aws.amazon.comgithub.comawslabs.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Hyperpod Performance Debugger loads about 4.1k tokens when it runs, and up to ~7.1k if it reads all its reference files. Until then it costs about 146 tokens; SKILL.md has 1,449 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from awslabs/agent-plugins at commit da51970, republished under its Apache-2.0 licence (© awslabs). 1,449 words, ~4,115 tokens.
.claude/skills/hyperpod-performance-debugger/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Route findings outside the two in-scope scenarios to the owner skill below.
| Concern observed | Route to |
|---|---|
| GPU hardware fault, ECC, NVLink, Xid, DCGM diagnostics, drain/replace | hyperpod-node-debugger (§ F Hardware/Auto-Repair, § G GPU) |
Cannot allocate memory at os.fork(), root volume exhausted | hyperpod-node-debugger (§ I Resource Exhaustion) |
| NCCL timeouts, hangs, AllReduce stalls, EFA TCP fallback, RDMA memlock | hyperpod-nccl |
| EFA / NCCL / CUDA / NVIDIA driver version drift across nodes | hyperpod-version-checker |
| EFA self-referencing security-group rule missing — single node | hyperpod-node-debugger § A (EFA / Security Group) |
| EFA self-referencing security-group rule missing — cluster-wide | hyperpod-cluster-debugger § A (EFA Health Checks) |
| Slurm node state changes (drain / resume / reboot) | hyperpod-slurm-debugger |
| Diagnostic bundle for AWS Support | hyperpod-issue-report |
| Shell access on a node | hyperpod-ssm |
hyperpod-version-checker.hyperpod-node-debugger § G.scripts/perf-snapshot.sh (read-only) to gather host-side signals for the suspect node and FSx filesystems mounted on it.[CONCERN] line in the script output, open the matching section below and read the supporting reference.bash scripts/perf-snapshot.sh --cluster <CLUSTER_NAME_OR_ARN> --region <REGION>
# Scope to one suspect node:
bash scripts/perf-snapshot.sh --cluster <C> --region <R> --node <INSTANCE_ID>The script samples one node by default. It collects host-side data via hyperpod-ssm: nvidia-smi output (temperature, SM clocks, PCIe link width, ECC, NVLink, topo -m), recent dmesg Xid lines, EFA port state and fi_info provider visibility, EFA installer + kernel module versions, CPU governor, NVL72 Fabric Manager state, FSx CloudWatch utilization, df -h / lfs df -h per mount, host iowait, /dev/shm size, and root-volume usage. All read-only.
Tags: [OK] healthy · [CONCERN] signal worth investigating (carries a → pointer to the owner skill) · [INFO] informational.
Host vs container scope. The script runs on the host via SSM and reports host-scope values. Many setups ship the EFA / libfabric / OFI-NCCL / CUDA stack inside the training container by design — a host value of unknown is not by itself a defect. What matters for performance is the stack the workload actually uses. Verify versions inside the container (and across nodes) via hyperpod-version-checker before drawing conclusions.
| Observation | Section |
|---|---|
| Pairwise NCCL bandwidth varies across node pairs / suspected straggler | A: Uneven NCCL Performance |
| Nodes spread across AZs / network-node-layer labels / UltraServer boundaries | A |
| EFA port not ACTIVE on a node, missing OFI plugin, or FI provider not visible | A + route to hyperpod-node-debugger § A; hyperpod-version-checker for cross-node version compare |
iostat shows high iowait, FSx CloudWatch utilization sustained near 100% | B: Poor Filesystem Performance |
| DataLoader stalls, checkpoint dominates step time | B |
Xid line in dmesg, uncorrectable ECC, inactive NVLink lane, GPU ≥ 88°C | Route to hyperpod-node-debugger § G |
| Container vs host version drift suspected | Route to hyperpod-version-checker |
Cannot allocate memory at os.fork(), root volume full, OOM events | Route to hyperpod-node-debugger § I |
NCCL timeout, hang, TCP fallback (NET/OFI Using TCP), RDMA memlock | Route to hyperpod-nccl |
The customer reports identical training jobs running with different step times on different node sets, pairwise bandwidth variance, or some allocations consistently slower than others despite identical code.
Per the official troubleshooting guide, the common contributing factors are network topology differences between nodes (cross-AZ, cross-rack, cross-UltraServer), degraded EFA performance on some nodes, mixed instance types or generations within an instance group, and CPU frequency scaling differences.
The host-side data points — GPU thermal/ECC/PCIe/clocks, Xid, NVLink lanes, EFA port state and provider visibility, CPU governor, EFA/OFI/driver versions, nvidia-smi topo -m — are all collected by scripts/perf-snapshot.sh (Step 1 above). The script tags [CONCERN] with thresholds and emits routing pointers; rerun it per suspect node via --node <INSTANCE_ID>.
For driver / CUDA / NCCL / EFA / OFI version drift across nodes, run hyperpod-version-checker skill.
Run the standard nccl-tests recipes from awslabs/awsome-distributed-training. For an N-node cluster, run all-reduce across every pair and record busbw for each pair. Pairs more than ~5% below the run mean (the threshold the AWS validation script flags) are problematic candidates.
Expected busbw per SKU is published in the AI-on-HyperPod NCCL test guide. Benchmark the specific instance type before relying on a number.
Pairwise scripts, HyperPod topology surfaces (HyperPod API, EKS labels, Slurm topology.conf), and GB200 NVL72 specifics are in references/perf-details.md § Uneven NCCL.
HyperPod exposes topology through three operator-visible surfaces:
aws sagemaker describe-cluster-node returns NodeDetails.Placement.AvailabilityZone / AvailabilityZoneId and NodeDetails.UltraServerInfo.Id (UltraServer SKUs only).topology.kubernetes.io/zone, topology.k8s.aws/network-node-layer-{1,2,3} (highest-numbered = closest to instance), topology.k8s.aws/ultraserver-id.topology.conf. Inspect via scontrol show topology.Tightly coupled work shares the same AZ, the same highest-numbered network-node-layer label (EKS) or the same Slurm topology block, and — for NVL72 jobs — the same UltraServerInfo.Id / topology.k8s.aws/ultraserver-id. If the cluster is spread across AZs or layers, topology must be re-established at provisioning time. Route provisioning changes to hyperpod-cluster-debugger § B (Capacity & AZ).
The customer reports training bottlenecked on data loading, checkpoint save/load dominating step time, executables/scripts loading slowly, or iowait high.
Per the official troubleshooting guide, the resolution path follows this order:
This skill covers steps 1–3. Steps 4–5 are customer decisions; surface the data and let the customer pick.
scripts/perf-snapshot.sh (Step 1 above) covers the on-node side of this pass: it discovers FSx mounts, calls aws cloudwatch get-metric-statistics on DataReadBytes and (for OpenZFS) FileServerDiskIopsUtilization, prints df -h for /fsx /opt/dlami/nvme /opt/sagemaker, runs lfs df -h per Lustre mount, and reports iostat iowait. It tags [CONCERN] when OpenZFS IOPS utilization sustains ≥ 80% or iowait > 20%.
For longer windows or additional metrics (DataWriteBytes, Lustre DiskIopsUtilization, OpenZFS FileServerDiskThroughputUtilization), drive the query directly:
aws cloudwatch get-metric-statistics --region <REGION> \
--namespace AWS/FSx --metric-name DataReadBytes \
--dimensions Name=FileSystemId,Value=<FSID> \
--start-time "$(date -u -d '3 hours ago' +%Y-%m-%dT%H:%M:%S)" \
--end-time "$(date -u +%Y-%m-%dT%H:%M:%S)" \
--period 60 --statistics Sum MaximumThe full per-filesystem-type metric catalog is in references/perf-details.md § Filesystem.
Provisioned capacity is saturated. CloudWatch utilization sustained near 100% across the workload window. Customer decision: scale up the filesystem.
StorageCapacity × PerUnitStorageThroughput; capacity changes are non-disruptive.I/O pattern is inefficient. CloudWatch shows headroom but the workload is still I/O-bound. Customer decision: change the application.
num_workers, set pin_memory=True, persistent_workers=True.torch.distributed.checkpoint.async_save plus FSDP SHARDED_STATE_DICT). FULL_STATE_DICT serializes through rank 0 and is a frequent root cause.Filesystem-selection guidance and the async-checkpoint pattern are in references/perf-details.md § Filesystem.
Once the immediate incident is diagnosed, recommend HyperPod's built-in health features so problems are caught before the next training run rather than after another customer-reported regression.
Enable NodeRecovery=Automatic on the cluster. The Health Monitoring Agent (HMA) continuously monitors GPU- and Trainium-based instances and marks instances unhealthy on detected failure. With auto-recovery enabled, HyperPod reboots or replaces the node — no operator intervention.
Enable OnStartDeepHealthChecks on every GPU instance group with both check categories:
InstanceStress — stress-ng on CPU/memory/disk, GPU and PCI device count verification, DCGM level-4 diagnostics (memory test included), and EFA loopback bandwidth/latency.InstanceConnectivity — multi-node NCCL all-reduce.Every newly provisioned or auto-replaced node passes the same hardware bar before accepting jobs.
Run on-demand deep health checks when this skill or any sibling surfaces a hardware concern but the cluster is mid-workload. aws sagemaker start-cluster-health-check runs the same checks against a specific instance group; nodes are placed in a Slurm maintenance reservation and the check is queued until any running job completes (not preempted). Console: HyperPod → Clusters → Instances → Run deep health checks.
Not supported when NodeProvisioningMode=Continuous; one on-demand request per cluster at a time. Requires the latest AMI — run UpdateClusterSoftware first.
Logs land in CloudWatch at /aws/sagemaker/Clusters/<cluster_name>/<cluster_id> under DeepHealthCheckResults/<log_stream_id>, and on each node at /var/log/aws/clusters/sagemaker-deep-health-check.log.
External:
busbw per SKU): https://awslabs.github.io/ai-on-sagemaker-hyperpod/docs/slurm-orchestration/validation-and-testing/performance-testing/nccl-tests© awslabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts, references) in plugins/sagemaker-ai/skills/hyperpod-performance-debugger of awslabs/agent-plugins.
Open the folder on GitHubat commit da51970
Hyperpod Performance Debugger next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Hyperpod Performance Debugger this skillawslabs/agent-plugins | 912 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Python Environment Setup for SageMakerhuggingface/skills | 11k | 2 repos | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Deployment Plannerhuggingface/skills | 11k | 1 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Hf Cloud Serving Image Selectionwaybarrios/opencode-power-pack | 533 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | |
| AWS AI MLaws/agent-toolkit-for-aws | 2.8k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.
huggingface/skills
Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.
waybarrios/opencode-power-pack
Select and verify the current region-specific serving container URI for a SageMaker model deployment.
aws/agent-toolkit-for-aws
Selects, deploys, and customizes AI models on Amazon SageMaker.
huggingface/skills
Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.
awslabs/agent-plugins
Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).
awslabs/agent-plugins
Generates code that transforms datasets between ML schemas for model training or evaluation.
awslabs/agent-plugins
Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes.
awslabs/agent-plugins
Evaluate, configure, and migrate workloads to AWS Lambda Managed Instances (LMI).
awslabs/agent-plugins
Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases.
awslabs/agent-plugins
Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters.
Works with
Categories
Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput. Hyperpod Performance Debugger is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.
Hyperpod Performance Debugger fits situations like: uneven NCCL across nodes; checkpoint slow; dataloader slow; filesystem bottleneck.
Run `npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a claude-code`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-performance-debugger in awslabs/agent-plugins) into .claude/skills/hyperpod-performance-debugger in your project. Claude Code loads it when a task matches its description.
Run `npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a codex`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-performance-debugger in awslabs/agent-plugins) into .agents/skills/hyperpod-performance-debugger in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awslabs/agent-plugins --skill hyperpod-performance-debugger -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hyperpod-performance-debugger, .gemini/skills/hyperpod-performance-debugger, .github/skills/hyperpod-performance-debugger and .opencode/skills/hyperpod-performance-debugger in your project.
Going by SKILL.md and its folder, Hyperpod Performance Debugger needs a shell for the scripts in its folder and the command-line tools its instructions call (aws and bash). Our summary lists: A Bash shell.
SKILL.md names 3 domains. As links in the text: docs.aws.amazon.com, github.com and awslabs.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Hyperpod Performance Debugger is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Hyperpod Performance Debugger: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Python Environment Setup for SageMaker (huggingface/skills, 11k stars), SageMaker Deployment Planner (huggingface/skills, 11k stars) and Hf Cloud Serving Image Selection (waybarrios/opencode-power-pack, 533 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
awslabs (a GitHub organization, an official publisher) maintains it in awslabs/agent-plugins, which has 912 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 5, 2026.
Source: awslabs/agent-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.