Aiml GPU Training Cluster Investigation
aws/tools-for-devops-agent
A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.
Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…
$ npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install awslabs/agent-plugins hyperpod-nccl --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-nccl .claude/skills/hyperpod-nccl && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "hyperpod-nccl" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-nccl into .claude/skills/hyperpod-nccl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-nccl", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-ncclType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install awslabs/agent-plugins hyperpod-nccl --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-nccl .agents/skills/hyperpod-nccl && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "hyperpod-nccl" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-nccl into .agents/skills/hyperpod-nccl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-nccl", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install awslabs/agent-plugins hyperpod-nccl --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-nccl .cursor/skills/hyperpod-nccl && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "hyperpod-nccl" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-nccl into .cursor/skills/hyperpod-nccl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-nccl", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/awslabs/agent-plugins.git --path plugins/sagemaker-ai/skills/hyperpod-nccl--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install awslabs/agent-plugins hyperpod-nccl --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-nccl .gemini/skills/hyperpod-nccl && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "hyperpod-nccl" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-nccl into .gemini/skills/hyperpod-nccl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-nccl", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install awslabs/agent-plugins hyperpod-ncclInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-nccl .github/skills/hyperpod-nccl && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "hyperpod-nccl" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-nccl into .github/skills/hyperpod-nccl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-nccl", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install awslabs/agent-plugins hyperpod-nccl --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-nccl .opencode/skills/hyperpod-nccl && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "hyperpod-nccl" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/hyperpod-nccl into .opencode/skills/hyperpod-nccl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperpod-nccl", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
hyperpod-ncclDiagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…
Hyperpod Nccl is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures, EFA TCP fallback, /dev/shm or memlock issues, NCCL version mismatch across pods, container OOM / exit-137 / OOMKilled, GPU OOM (CUDA out of memory), CrashLoopBackOff / Pending pods, MASTERADDR DNS, NetworkPolicy blocking. Not for single-node hardware faults (→ hyperpod-node-debugger § G) or cluster-creation EFA / SSM…
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/debugging-guide.md`, `references/error-patterns-quick-ref.md` and `references/operations.md`).
It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with Amazon SageMaker, Amazon Web Services and CUDA. The repository describes itself as: Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS. The licence is Apache-2.0.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit da51970. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
awsbashkubectlyumaptFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use aws and kubectl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Hyperpod Nccl loads about 3.4k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 145 tokens; SKILL.md has 984 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from awslabs/agent-plugins at commit da51970, republished under its Apache-2.0 licence (© awslabs). 984 words, ~3,420 tokens.
.claude/skills/hyperpod-nccl/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Operating policy. Run read-only diagnostics yourself. Never run a command that changes cluster, node, or workload state — present each one as a Suggested command (run this yourself) block and wait for the customer. Destructive order: investigate → reboot → replace (replace destroys root + secondary volumes; not supported on Slurm controller nodes). Never discard training state on speculation.
Diagnose NCCL failures on SageMaker HyperPod (EKS and Slurm). scripts/nccl-diagnose.sh reads state via AWS APIs, kubectl, and SSM, then prints each issue as [FAIL] ... → references/<file>.md § <section>. Read-only.
Signal sourcing: list-cluster-events carries infrastructure-level state only (lifecycle, bootstrap, EFA health check, capacity, replacement, reboot, AMI rollback). It does not carry NCCL timeouts, GPU XID/ECC, or per-pod training signals — those come from pod logs, CloudWatch training streams, on-node SSM probes, and NCCL env audit. "No events" on a training-time NCCL issue is expected, not a clean bill of health.
[FAIL] line, Read the referenced section.If a finding has no matching section, report it as a bug — do not invent a fix.
EKS_ARN=$(aws sagemaker describe-cluster --cluster-name <HYPERPOD-NAME> --region <REGION> \
--query 'Orchestrator.Eks.ClusterArn' --output text)
EKS_NAME=$(echo "$EKS_ARN" | awk -F'/' '{print $NF}')
aws eks update-kubeconfig --name "$EKS_NAME" --region <REGION>
kubectl get nodes# Basic:
bash scripts/nccl-diagnose.sh --cluster <HYPERPOD-NAME> --region <REGION>
# Scope to an EKS job/namespace:
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --namespace <NS> --job <JOB>
# Force orchestrator:
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --orchestrator slurm
# Larger hardware sample (default 3):
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --sample-nodes 10
# Specific node only:
bash scripts/nccl-diagnose.sh --cluster <NAME> --region <REGION> --node i-0abc123def456Tags: [PASS] · [FAIL] (counted in Issues Found, has reference pointer) · [WARN] · [INFO]. Priorities: P0 blocks training · P1 degraded · P2 informational.
Each [FAIL] line in the script already points directly at the right section. This table is a lookup for manual triage.
| Finding | Section |
|---|---|
| SG missing inbound/outbound self-reference | operations.md § 8 |
| Blocking NetworkPolicy / allow-all missing | operations.md § 8 |
| Slurm node DOWN / DRAINING / RemoveIPC | operations.md § 7 |
| GPU XID / SYSTEM_ERROR / hardware fault | hyperpod-node-debugger § F / § G |
| GPU row-remap / DCGM Fail / silent NaNs | hyperpod-node-debugger § G.1.a/b |
| NCCL timeout / rendezvous / straggler | debugging-guide.md § 1 |
| EFA configuration / not used | debugging-guide.md § 6 |
EFA TCP fallback (NET/OFI Using TCP) | debugging-guide.md § 13 |
| NCCL version mismatch across pods | debugging-guide.md § 10 |
| Container OOM (pod killed, exit 137) | debugging-guide.md § 4 |
GPU OOM (CUDA out of memory) | debugging-guide.md § 11 |
RDMA memlock / /dev/shm too small | debugging-guide.md § 17 |
| MASTER_ADDR DNS / headless Service | debugging-guide.md § 12 |
| NVLS / PXN / topology tuning | debugging-guide.md § 19 |
| Any NCCL / EFA / rendezvous log pattern | error-patterns-quick-ref.md |
| Performance / nccl-tests / bandwidth | performance-testing.md |
aws CLI v2.13+ authenticated (aws sts get-caller-identity)jq, python3, bash 4.2+unbuffer (from the expect package: yum install expect / apt install expect)kubectl authenticated to the EKS cluster (K8s checks skipped if absent)session-manager-plugin for on-node hardware checks--region or set $AWS_DEFAULT_REGION.--orchestrator eks|slurm.--namespace <NS> --job <JOB>.--node <ID> for a specific node. Node probes run serially (180 s per node): --sample-nodes 10 can take ~30 min.TERM=dumb.| Failure | Script | Tell the customer |
|---|---|---|
aws sts get-caller-identity fails | Exit 1 with the AWS error | "Fix AWS credentials and rerun." |
describe-cluster AccessDenied | Warn, add Missing IAM for sagemaker:DescribeCluster | "Grant sagemaker:DescribeCluster (operations.md § 2)." |
| Cluster not found | Exit 1 after listing region's clusters | "Confirm HyperPod cluster name and region." |
kubectl absent / unauthenticated | Warn, skip K8s checks | "aws eks update-kubeconfig --name <EKS> --region <R>." |
| SSM plugin absent | Warn, skip on-node hardware checks | "Install session-manager-plugin." |
| SSM times out (180s) | Partial output, mark node unreachable | "Rerun with --node <ID> --sample-nodes 1; check SSM agent on the node." |
| CloudWatch log group not found | Skip CloudWatch scan | "Enable CloudWatch on the cluster (operations.md § 4)." |
| Cluster events API throttled | Warn, continue with partial data | "Rerun later — script is idempotent." |
Exit codes: 0 diagnostic complete · 1 fatal prerequisite missing or cluster unreachable.
Full policy + RBAC in operations.md § 2. SSM on HyperPod uses start-session against sagemaker-cluster:<cluster-id>_<group>-<iid> targets — grant ssm:StartSession / ssm:TerminateSession, not ssm:SendCommand.
| Scope | Method | Coverage |
|---|---|---|
| All nodes | sagemaker:ListClusterNodes (paginated) | 100% nodes |
| All K8s objects | kubectl | 100% pods/nodes/policies |
| Hardware | SSM --sample-nodes N (default 3) | Sampled |
| Node logs | CloudWatch | 100% nodes |
Large clusters: the PyTorch NCCL backend defaults to a 10-minute collective-op timeout (per the PyTorch distributed docs). Large clusters routinely exceed that on first rendezvous; raise it via torch.distributed.init_process_group(timeout=timedelta(seconds=<N>)). HyperPod support has also observed NCCL topology-graph-search hangs on 256+ node clusters when memlock is unlimited; using a large fixed memlock (e.g. 8388608) in pod securityContext or /etc/security/limits.conf has cleared these in field cases. This memlock pattern is a field observation, not AWS- or NCCL-documented behavior.
For FSDP, DeepSpeed, or Megatron-LM tuning: debugging-guide.md § 18.
| Need | Use |
|---|---|
| Cluster creation / deployment failures | hyperpod-cluster-debugger (§ A / B / C / H + --validate) |
| Post-deployment cluster-wide management | hyperpod-cluster-debugger |
| Per-node issues (disk, lifecycle, hardware) | hyperpod-node-debugger |
| Trainium/Inferentia collective-comm (AWS Neuron Collectives, not NCCL) | hyperpod-node-debugger § G.2 |
| Shell on nodes | hyperpod-ssm |
| Version comparison across nodes | hyperpod-version-checker |
| Diagnostic bundle for AWS Support | hyperpod-issue-report |
| MFU / performance degradation | hyperpod-mfu-debugger |
Escalate when:
Issues Found: 0 but training still fails.# 1. Cluster identity + status
aws sagemaker describe-cluster --cluster-name <C> --region <R>
# 2. Full NCCL diagnostic (sample more nodes for escalation)
bash scripts/nccl-diagnose.sh --cluster <C> --region <R> --sample-nodes 10 > nccl-diag.txt
# 3. Per-node log/config bundle to S3 (delegates to hyperpod-issue-report)
# See skills/hyperpod-issue-report/SKILL.md for the exact invocation.nccl-diag.txt from step 2 abovehyperpod-issue-report bundle from step 3printenv | grep -E '^NCCL|^FI_|^TORCH_' from one pod)© awslabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in plugins/sagemaker-ai/skills/hyperpod-nccl of awslabs/agent-plugins.
Open the folder on GitHubat commit da51970
Hyperpod Nccl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Hyperpod Nccl this skillawslabs/agent-plugins | 915 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | |
| Aiml GPU Training Cluster Investigationaws/tools-for-devops-agent | 100 | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Paddle Design CompilerPaddlePaddle/Paddle | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| Python Environment Setup for SageMakerhuggingface/skills | 11k | 2 repos | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Index Widthpytorch/pytorch | 104k | — | ~1.6k | Automated safety check: Pass | Custom licence |
aws/tools-for-devops-agent
A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
PaddlePaddle/Paddle
A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…
huggingface/skills
Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
awslabs/agent-plugins
Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).
awslabs/agent-plugins
Generates code that transforms datasets between ML schemas for model training or evaluation.
awslabs/agent-plugins
Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes.
awslabs/agent-plugins
Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases.
awslabs/agent-plugins
Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.
awslabs/agent-plugins
Remote command execution and file transfer on SageMaker HyperPod cluster nodes via AWS Systems Manager (SSM).
Works with
Categories
Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…. Hyperpod Nccl is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures, EFA TCP fallback, /dev/shm or memlock issues, NCCL version mismatch across pods, container OOM / exit-137 / OOMKilled, GPU OOM (CUDA out of memory), CrashLoopBackOff / Pending pods, MASTERADDR DNS, NetworkPolicy blocking.
Hyperpod Nccl fits situations like: tasks that involve GPU and accelerator computing.
Run `npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a claude-code`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-nccl in awslabs/agent-plugins) into .claude/skills/hyperpod-nccl in your project. Claude Code loads it when a task matches its description.
Run `npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a codex`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-nccl in awslabs/agent-plugins) into .agents/skills/hyperpod-nccl in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awslabs/agent-plugins --skill hyperpod-nccl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hyperpod-nccl, .gemini/skills/hyperpod-nccl, .github/skills/hyperpod-nccl and .opencode/skills/hyperpod-nccl in your project.
Going by SKILL.md and its folder, Hyperpod Nccl needs a shell for the scripts in its folder and the command-line tools its instructions call (aws, bash, kubectl, yum and apt). Our summary lists: A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Hyperpod Nccl is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 22k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Hyperpod Nccl: Aiml GPU Training Cluster Investigation (aws/tools-for-devops-agent, 100 stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars) and Python Environment Setup for SageMaker (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
awslabs (a GitHub organization, an official publisher) maintains it in awslabs/agent-plugins, which has 915 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 5, 2026.
Source: awslabs/agent-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.