Official agent skill

Hyperpod Cluster Debugger

by awslabs in awslabs/agent-plugins

Diagnose and remediate cluster-wide HyperPod (EKS or Slurm) problems — creation / deployment failures (CloudFormation, EFA health check, lifecycle scripts, capacity), EKS access, node replacement…

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Hyperpod Cluster Debugger

skills CLI
$ npx skills add awslabs/agent-plugins --skill hyperpod-cluster-debugger -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install awslabs/agent-plugins hyperpod-cluster-debugger --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sagemaker-ai/skills/hyperpod-cluster-debugger .claude/skills/hyperpod-cluster-debugger && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hyperpod-cluster-debugger
GitHub stars
915
Used in
1 other repo
Token cost
~3.6k tokens
SKILL.md length
1,208 words
Files
8 (incl. scripts, references)
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnose and remediate cluster-wide HyperPod (EKS or Slurm) problems — creation / deployment failures (CloudFormation, EFA health check, lifecycle scripts, capacity), EKS access, node replacement…

  • Works in 2 steps: Run diagnostics → Match signal → section
  • Tasks that involve Infrastructure as code
  • SKILL.md covers Workflow, Step 1: Run diagnostics, Step 2: Match signal → section and A: EFA Health Checks, plus 16 more sections
  • Runs Shell scripts from its folder; calls bash, aws and kubectl

What it does

Hyperpod Cluster Debugger is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose and remediate cluster-wide HyperPod (EKS or Slurm) problems — creation / deployment failures (CloudFormation, EFA health check, lifecycle scripts, capacity), EKS access, node replacement, CloudFormation nested-stack errors, post-maintenance rollback state, dangling nodes, autoscaler conflicts. Includes --validate pre-flight. Read-only.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/capacity-planning.md`, `references/cloudformation-errors.md` and `references/cluster-diagnostics-detail.md`).

It sits in DevOps & Cloud, covering Infrastructure as code. It works with AWS CloudFormation and Amazon Web Services. The repository describes itself as: Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Infrastructure as code

Example prompts

  • “/hyperpod-cluster-debugger”

Requirements

  • A Bash shell

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Run diagnostics
  2. Match signal → section

What it can do on your machine

Read from SKILL.md and the folder at commit da51970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • aws
    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use aws and kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hyperpod Cluster Debugger loads about 3.6k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 94 tokens; SKILL.md has 1,208 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~17k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from awslabs/agent-plugins at commit da51970, republished under its Apache-2.0 licence (© awslabs). 1,208 words, ~3,604 tokens.

Download SKILL.mdSave it as .claude/skills/hyperpod-cluster-debugger/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
hyperpod-cluster-debugger
description
Diagnose and remediate cluster-wide HyperPod (EKS or Slurm) problems — creation / deployment failures (CloudFormation, EFA health check, lifecycle scripts, capacity), EKS access, node replacement, CloudFormation nested-stack errors, post-maintenance rollback state, dangling nodes, autoscaler conflicts. Includes `--validate` pre-flight. Read-only.
metadata.version
0.0.1

HyperPod Cluster Debugger

Operating policy. Run read-only diagnostics yourself. Never run a command that changes cluster, node, or workload state — present each one as a Suggested command (run this yourself) block and wait for the customer to run it. Destructive order: investigate → reboot → replace (replace destroys root + secondary volumes; not supported on Slurm controller nodes).

Before any state-changing CLI: ask if it's IaC-managed. HyperPod clusters, SGs, EKS access entries, and IAM are usually provisioned via CloudFormation / CDK / Terraform. If yes, the fix belongs in IaC — running the CLI will drift and the next deploy reverts it. Use the CLI only when IaC is unavailable (locked out, predates IaC, mid-review).

scripts/diagnose-cluster.sh is read-only: it collects state via AWS APIs (and SSM for Slurm controller health) and prints each issue as [FAIL] ... → references/<file>.md § <section>.

ReferenceOpen when
cluster-diagnostics-detail.mdPer-finding remediation runbook (§ A–L)
cluster-operations.mdOperational deep-dives (EFA SG, EKS access, SSM, Slurm, filesystem)
cloudformation-errors.md§ H needs the full per-resource CFN error catalog
capacity-planning.md§ B or --validate flags capacity / subnet sizing
lifecycle-scripts.md§ C points at a specific lifecycle failure
iam-permissions.mdFull IAM policy for the diagnostic

Workflow

  1. Collect HyperPod cluster name (not EKS name), region, exact error string.
  2. Run scripts/diagnose-cluster.sh (or --validate for pre-create).
  3. For every [FAIL] line, Read the referenced section.
  4. Present finding, root cause, and the Suggested-command block verbatim. Wait for customer approval.
  5. Re-run the diagnostic to confirm.

Step 1: Run diagnostics

bash
# Diagnose an existing cluster:
bash scripts/diagnose-cluster.sh --cluster <CLUSTER_NAME_OR_ARN> --region <REGION>

# Pre-flight (no cluster needed) — validates SGs, subnets, IAM, VPC endpoints,
# optionally S3 lifecycle scripts and per-AZ capacity:
bash scripts/diagnose-cluster.sh --validate --region <REGION> \
  --sg-ids <sg-1,sg-2> --subnet-ids <sub-1,sub-2> [--iam-role <role-arn>] \
  [--s3-uri s3://<BUCKET>/path/] [--instance-type ml.p5.48xlarge]

Pass --instance-type when the target instance type is known — enables the per-AZ capacity check (warns if none of the provided subnets are in an AZ that offers that type, which causes insufficient-capacity failures at creation time).

Tags: [PASS] · [FAIL] (counted, has → references/... pointer) · [WARN] · [INFO]. Priorities: P0 blocks operation · P1 degraded · P2 informational.


Step 2: Match signal → section

Error messages / events:

SignalSection
"EFA health checks did not run successfully" (public-doc verbatim signal)A: EFA Health Checks
Insufficient-capacity or AZ-mismatch failure at creationB: Capacity & AZ
Lifecycle-script failure or timeout during provisioningC: Lifecycle Scripts
kubectl auth error (server asks for credentials / no API group list)D: EKS Access
InService but not all instances visibleE: Cluster Provisioning
"Target is not connected" / SSM errorsF: SSM Connectivity
Node replacement not happening / batch-replace not workingG: Node Replacement
"Embedded stack failed" / any CloudFormation errorH: CloudFormation Errors
UpdateClusterSoftware failed or cluster in post-maintenance rollback stateJ: AMI & Cluster Updates
Dangling / orphaned nodes in EKS vs list-cluster-nodesK: Dangling Nodes & Cleanup
Cluster Autoscaler breaks after HyperPod attachedL: Autoscaler Compatibility
Slow I/O, FSx throughput saturatedcluster-operations.md § 9
Slurm node name → instance ID lookupI: Utilities

A: EFA Health Checks

SG missing self-reference. Add inbound + outbound self-ref to every SG on the cluster, plus least-privilege egress for the AWS APIs the node needs (HTTPS 443 to S3 / ECR / SageMaker / SSM / STS / CloudWatch Logs — via VPC-endpoint prefix-lists when possible). Full procedure: cluster-diagnostics-detail.md § A.

B: Capacity & AZ

Instance type unavailable in the requested AZ. Verify with describe-instance-type-offerings, then change AZ, use Flexible Training Plans, or request ODCR. Full: § B · strategy: capacity-planning.md.

C: Lifecycle Scripts

Script failed or timed out during provisioning. Read CloudWatch under /aws/sagemaker/Clusters/<name>/<id> — common causes: missing S3 VPC endpoint, IAM gap, CRLF line endings, instance-group name mismatch. Full: § C · layout: lifecycle-scripts.md.

D: EKS Access / kubectl

IAM identity not in EKS access entries. Verify with sts get-caller-identity, create an access entry with admin policy, update kubeconfig. Full: § D.

E: Cluster Provisioning

InService without all instances is expected under Continuous Provisioning — failures surface as events, not cluster errors. For stuck Creating/Updating/Deleting: check CFN nested stacks (§ H), IAM, capacity, events; if stuck Deleting check VPC ENI dependencies. Full: § E.

F: SSM Connectivity

Target is not connected: use sagemaker-cluster:<CLUSTER_ID>_<GROUP>-<INSTANCE_ID> format (not raw EC2 ID), install session-manager-plugin, confirm node Running. Check IAM + VPC endpoints on timeouts. Full: § F.

G: Node Replacement

Auto-repair: confirm NodeRecovery=Automatic, check Health Monitoring Agent (HMA) logs + node labels / Slurm reason, confirm capacity. Manual: reboot first, replace only if reboot fails. Replace requires the cluster to have been patched via UpdateClusterSoftware at least once and cannot target a Slurm controller node. Full: § G.

H: CloudFormation Errors

Embedded stack failed hides the real error. Drill into nested stacks via Events tab (filter Failed) until you reach a non-stack resource. CLI: describe-stack-events --query 'StackEvents[?ResourceStatus==\CREATE_FAILED`]'`. Also covers SLR creation failures and permission-boundary denials. Full: § H · catalog: cloudformation-errors.md.

I: Utilities

Map Slurm node names (ip-10-x-y-z) to HyperPod instance IDs via list-cluster-nodes or on-node /opt/ml/config/resource_config.json. Full: § I.

Show full SKILL.md (476 more words)Show less

J: AMI & Cluster Updates

UpdateClusterSoftware fails and rolls back, or the cluster stays in a post-maintenance rollback state. Common causes: lifecycle script incompatible with new AMI, HMA version too old, insufficient rolling-update capacity. If the cluster has active nodes, collect diagnostics and escalate rather than delete-and-recreate. Full: § J.

K: Dangling Nodes & Cleanup

Nodes in kubectl get nodes but not in list-cluster-nodes (ghost EKS nodes), or the inverse (HyperPod nodes that never registered kubelet). Script flags both. Full: § K.

L: Autoscaler Compatibility

Cluster Autoscaler errors on HyperPod provider IDs and breaks autoscaling for all node groups. No officially endorsed workaround — escalate to AWS Support. Karpenter does not conflict with HyperPod nodes by default. Full: § L.


Prerequisites

  • aws CLI v2.13+ authenticated to the cluster's account
  • jq, python3, bash 4.2+
  • kubectl authenticated to the EKS cluster (EKS checks skipped if absent)
  • session-manager-plugin (Slurm controller health checks only)

IAM policy: references/iam-permissions.md.

Defaults

  • Region — required: pass --region or set $AWS_DEFAULT_REGION.
  • Mode — --cluster <NAME> (diagnose) or --validate (pre-create).
  • Event window — up to 500 most recent events (5 × 100, paginated).
  • Colors — auto-disabled on non-TTY; --no-color to force off.

Error handling

FailureScriptTell the customer
aws sts get-caller-identity failsExit 1"Fix AWS credentials and rerun."
Cluster not foundExit 1 after listing region's clusters"Confirm HyperPod cluster name (not EKS) and region."
sagemaker:* / ec2:* / eks:* / logs:* deniedWarn, add Missing IAM permission for <API>, continue"Grant the listed IAM action and rerun."
kubectl absent or unauthenticatedSkip EKS checks (access entries, add-ons, aws-auth, nodes)"Install/authenticate kubectl."
session-manager-plugin absent (Slurm)Skip Slurm controller probe"Install session-manager-plugin."
SSM throttled / times out (180s)Retry with backoff; warn and continue if still failing"Rerun later — script is idempotent."
CloudWatch log group not foundSkip CloudWatch check"CloudWatch not configured on this cluster."

Exit codes: 0 no critical failures · 1 one or more critical failures (cluster not found, fatal prerequisite missing, or any [FAIL] in diagnose or --validate mode). [WARN] lines do not affect the exit code.

Skill delegation

NeedUse
Shell on nodeshyperpod-ssm
Version comparison across nodeshyperpod-version-checker

Escalate to AWS Support

Escalate when:

  1. EFA health checks fail despite correct SG rules.
  2. Capacity errors persist despite a valid Flexible Training Plan / ODCR.
  3. Node replacement fails repeatedly without clear events / log signal.
  4. Cluster stuck in a non-terminal state (Creating, Updating, or a post-maintenance rollback state) for an extended period.
  5. CloudFormation root-cause is an internal service error.
Before opening the case

Run these commands and attach the output. Goal: AWS Support has everything at case open.

bash
# 1. Cluster identity + status (confirms region, ARN, orchestrator, instance groups)
aws sagemaker describe-cluster --cluster-name <CLUSTER> --region <REGION>

# 2. Full cluster-level diagnostic bundle
bash scripts/diagnose-cluster.sh --cluster <CLUSTER> --region <REGION> > diag.txt

# 3. Per-node log/config bundle to S3 (delegates to hyperpod-issue-report skill)
#    See skills/hyperpod-issue-report/SKILL.md for the exact invocation.
Include in the case
  • Cluster name + ARN (or ClusterId suffix) and AWS region
  • ClusterStatus + FailureMessage from describe-cluster
  • Timestamp window (UTC start / end) of the failure
  • Exact error strings observed (copy verbatim from events / logs / console)
  • Affected instance IDs / NodeLogicalIds / instance group names
  • diag.txt from step 2 above
  • S3 URI of the hyperpod-issue-report bundle from step 3

© awslabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in plugins/sagemaker-ai/skills/hyperpod-cluster-debugger of awslabs/agent-plugins.

  • SKILL.md
  • references/capacity-planning.md
  • references/cloudformation-errors.md
  • references/cluster-diagnostics-detail.md
  • references/cluster-operations.md
  • references/iam-permissions.md
  • references/lifecycle-scripts.md
  • scripts/diagnose-cluster.sh

Open the folder on GitHubat commit da51970

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in awslabs/agent-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Hyperpod Cluster Debugger next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hyperpod Cluster Debugger compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hyperpod Cluster Debugger this skillawslabs/agent-plugins9151 repos~3.6kAutomated safety check: PassApache-2.0
AWS Cdk Developmentzxkane/aws-skills3672 repos~2.5kAutomated safety check: PassMIT
AWS Cloud Advisortech-leads-club/agent-skills7k—~2.1kAutomated safety check: PassCC-BY-4.0
AWS Native Runtime Investigationpulumi/pulumi-aws-native108—~753Automated safety check: PassApache-2.0
Hunt Bugsgo-to-k/cdkd143—~2.2kAutomated safety check: PassApache-2.0
Cloudformationitsmostafa/aws-agent-skills1.2k—~2.5kAutomated safety check: PassMIT

Similar skills

  • AWS Cdk Development

    zxkane/aws-skills

    AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.

    367 GitHub starsUsed in 2 repos~2.5k tokens
    DevOps & CloudAuto-check passed
  • AWS Cloud Advisor

    tech-leads-club/agent-skills

    Answers AWS architecture, security and service-selection questions by searching AWS documentation through MCP tools first, then adapting advice to your stack and team.

    7k GitHub stars~2.1k tokensUpdated 18 days ago
    DevOps & CloudAuto-check passed
  • AWS Native Runtime Investigation

    pulumi/pulumi-aws-native

    Official

    Use after triage or repository evidence establishes that an issue involves Pulumi AWS Native runtime behavior across the Pulumi provider protocol, generated CloudFormation metadata, and AWS Cloud…

    108 GitHub stars~753 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Hunt Bugs

    go-to-k/cdkd

    Proactively hunt for cdkd bugs by deploying real CDK apps that exercise common-but-untested AWS resources, configs, and CloudFormation notations against real AWS, then fix what breaks.

    143 GitHub stars~2.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Cloudformation

    itsmostafa/aws-agent-skills

    AWS CloudFormation infrastructure as code for stack management.

    1.2k GitHub stars~2.5k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Cdkd

    go-to-k/cdkd

    Install cdkd and use it safely from an AWS CDK project. An agent skill from go-to-k/cdkd.

    143 GitHub stars~6.7k tokensUpdated today
    DevOps & CloudAuto-check: notes

More from awslabs/agent-plugins

All 33 skills in this repo
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    915 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    915 GitHub starsUsed in 2 repos~3.5k tokens
    Auto-check passed
  • Finetuning Technique

    awslabs/agent-plugins

    Official

    Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes.

    915 GitHub starsUsed in 1 repo~604 tokens
    Auto-check passed
  • Hyperpod Issue Report

    awslabs/agent-plugins

    Official

    Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases.

    915 GitHub starsUsed in 1 repo~890 tokens
    Auto-check passed
  • Hyperpod Performance Debugger

    awslabs/agent-plugins

    Official

    Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

    915 GitHub starsUsed in 1 repo~4.1k tokens
    Auto-check passed
  • Hyperpod Ssm

    awslabs/agent-plugins

    Official

    Remote command execution and file transfer on SageMaker HyperPod cluster nodes via AWS Systems Manager (SSM).

    915 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check: notes

Categories

Questions about Hyperpod Cluster Debugger

What does Hyperpod Cluster Debugger do?

Diagnose and remediate cluster-wide HyperPod (EKS or Slurm) problems — creation / deployment failures (CloudFormation, EFA health check, lifecycle scripts, capacity), EKS access, node replacement…. Hyperpod Cluster Debugger is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Diagnose and remediate cluster-wide HyperPod (EKS or Slurm) problems — creation / deployment failures (CloudFormation, EFA health check, lifecycle scripts, capacity), EKS access, node replacement, CloudFormation nested-stack errors, post-maintenance rollback state, dangling nodes, autoscaler conflicts.

When should I use Hyperpod Cluster Debugger?

Hyperpod Cluster Debugger fits situations like: tasks that involve Infrastructure as code.

How do I install Hyperpod Cluster Debugger in Claude Code?

Run `npx skills add awslabs/agent-plugins --skill hyperpod-cluster-debugger -a claude-code`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-cluster-debugger in awslabs/agent-plugins) into .claude/skills/hyperpod-cluster-debugger in your project. Claude Code loads it when a task matches its description.

How do I install Hyperpod Cluster Debugger in Codex?

Run `npx skills add awslabs/agent-plugins --skill hyperpod-cluster-debugger -a codex`. Or copy the skill folder (plugins/sagemaker-ai/skills/hyperpod-cluster-debugger in awslabs/agent-plugins) into .agents/skills/hyperpod-cluster-debugger in your project. Codex loads it when a task matches its description.

Can I use Hyperpod Cluster Debugger in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awslabs/agent-plugins --skill hyperpod-cluster-debugger -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hyperpod-cluster-debugger, .gemini/skills/hyperpod-cluster-debugger, .github/skills/hyperpod-cluster-debugger and .opencode/skills/hyperpod-cluster-debugger in your project.

What does Hyperpod Cluster Debugger need to run?

Going by SKILL.md and its folder, Hyperpod Cluster Debugger needs a shell for the scripts in its folder and the command-line tools its instructions call (bash, aws and kubectl). Our summary lists: A Bash shell.

Does Hyperpod Cluster Debugger access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Hyperpod Cluster Debugger safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Hyperpod Cluster Debugger use?

Hyperpod Cluster Debugger is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hyperpod Cluster Debugger use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.

What are the alternatives to Hyperpod Cluster Debugger?

Skills that share tags, products or a category with Hyperpod Cluster Debugger: AWS Cdk Development (zxkane/aws-skills, 367 stars), AWS Cloud Advisor (tech-leads-club/agent-skills, 7k stars), AWS Native Runtime Investigation (pulumi/pulumi-aws-native, 108 stars) and Hunt Bugs (go-to-k/cdkd, 143 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hyperpod Cluster Debugger?

awslabs (a GitHub organization, an official publisher) maintains it in awslabs/agent-plugins, which has 915 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 5, 2026.

Source: awslabs/agent-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.