Search

Amazon Web Services · GPU and accelerator computing

5 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost.

Orchestra-Research/AI-Research-SKILLs13k4 repos~2.4kAutomated safety check: PassMIT3 mo ago
2

A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.

aws/tools-for-devops-agent103—~5.4kAutomated safety check: PassApache-2.0yesterday
3

Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

awslabs/agent-plugins916—~4.1kAutomated safety check: PassApache-2.0yesterday
4

Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

awslabs/agent-plugins916—~910Automated safety check: PassApache-2.0yesterday
5
5.Hyperpod NcclOfficial

Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…

awslabs/agent-plugins916—~3.4kAutomated safety check: PassApache-2.0yesterday