Topic · AI & LLM Engineering

Best GPU and accelerator computing skills, page 2

Skills #49–96 of 176, ranked by score.

GPU and accelerator computing skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

GPU and accelerator computing skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.4kAutomated safety check: PassMIT3 mo ago
50

Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits.

huggingface/skills11k2 repos~4.6kAutomated safety check: PassApache-2.06 days ago
51

Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

BBuf/AI-Infra-Auto-Driven-SKILLS911—~2kAutomated safety check: PassNo licence2 days ago
52

Organizes PyTorch training code into LightningModules, DataModules and Trainers, with multi-GPU strategies, callbacks and logging configured.

davila7/claude-code-templates32k14 repos~1.7kAutomated safety check: PassMITtoday
53

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

slowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT2 mo ago
54

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
55

Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI.

lucifer1004/VeloQ127—~1.3kAutomated safety check: PassMIT3 days ago
56
56.Cuda

Draft, debug, and measure CUDA kernels and host launch workflows.

sablin39/tilelang-cuda-skills145—~990Automated safety check: PassNo licence22 days ago
57

Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably.

Krusty84/triton-ascend-agent-dev-kit106—~564Automated safety check: PassApache-2.01 mo ago
58

Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing.

microsoft/onnxruntime22k—~6.5kAutomated safety check: PassMITtoday
59
59.At Dispatch V2Official

Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

intel/torch-xpu-ops1153 repos~2.2kAutomated safety check: PassApache-2.0today
60

PyTorch training reference: architecture choice by data type, scaling rules, a training loop, optimizer and learning-rate choices, and fixes for loss spikes or OOM.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.8kAutomated safety check: PassMIT3 mo ago
61

Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use.

davila7/claude-code-templates32k10 repos~2.4kAutomated safety check: PassMITtoday
62

Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

Orchestra-Research/AI-Research-SKILLs13k3 repos~1.8kAutomated safety check: PassMIT3 mo ago
63

A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.

aws/tools-for-devops-agent100—~5.4kAutomated safety check: PassApache-2.0today
64

Review 3DGS implementation code for correctness, performance bugs, and best practices.

jaccen/Awesome-Gaussian-Skills161—~2.9kAutomated safety check: PassApache-2.0today
65

Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.

Krusty84/triton-ascend-agent-dev-kit106—~597Automated safety check: PassApache-2.01 mo ago
66

Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.

Orchestra-Research/AI-Research-SKILLs13k5 repos~3kAutomated safety check: WarnMIT3 mo ago
67

Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

BBuf/AI-Infra-Auto-Driven-SKILLS911—~3.9kAutomated safety check: PassNo licence2 days ago
68

Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

NVIDIA/skills3.5k—~4.7kAutomated safety check: NotesApache-2.0today
69

Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

Orchestra-Research/AI-Research-SKILLs13k1 repo~3.7kAutomated safety check: PassMIT3 mo ago
70

Fine-tunes and serves Physical Intelligence's pi0, pi0-fast and pi0.5 robot policies with JAX or PyTorch, including checkpoint conversion and policy servers.

Orchestra-Research/AI-Research-SKILLs13k1 repo~3.6kAutomated safety check: PassMIT3 mo ago
71

Write optimized Triton GPU kernels for deep learning operations.

vipshop/cache-dit1.3k—~1.1kAutomated safety check: PassApache-2.09 days ago
72

Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

awslabs/agent-plugins9151 repo~910Automated safety check: PassApache-2.02 days ago
73

Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

BBuf/AI-Infra-Auto-Driven-SKILLS911—~7.5kAutomated safety check: PassNo licence2 days ago
74

Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

facebookexperimental/triton201—~709Automated safety check: PassMITtoday
75

Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs.

Krusty84/triton-ascend-agent-dev-kit106—~662Automated safety check: PassApache-2.01 mo ago
76

Guided step-by-step wizard for producing a new FlyDSL GPU kernel from a requirement: classify the kernel type, pick a skeleton, fill in compute, add control flow / sync / LDS, then test on GPU.

ROCm/FlyDSL288—~4.6kAutomated safety check: NotesUnknowntoday
77

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills398—~2.3kAutomated safety check: PassMITtoday
78
78.Ir DebuggingOfficial

Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

facebookexperimental/triton201—~644Automated safety check: PassMITtoday
79

Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation.

Krusty84/triton-ascend-agent-dev-kit106—~753Automated safety check: PassApache-2.01 mo ago
80

Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling.

Orchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT3 mo ago
81

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups

sgl-project/sglang37k—~13kAutomated safety check: PassApache-2.0today
82

Estimate peak GPU memory for a DualPipeV training run. An agent skill from mlc-ai/pith-train.

mlc-ai/pith-train355—~1.9kAutomated safety check: PassApache-2.03 days ago
83

3DGS compression-to-deployment pipeline: quantization (scalar/VQ/mixed-precision), pruning (coreset/adaptive/variational/merge/Bayesian), progressive streaming & LoD, Web/WebGPU/mobile deployment…

jaccen/Awesome-Gaussian-Skills161—~5.5kAutomated safety check: PassApache-2.0today
84

Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs.

Krusty84/triton-ascend-agent-dev-kit106—~636Automated safety check: PassApache-2.01 mo ago
85

Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs.

ZJLi2013/awesome-kernel-skills102—~1.1kAutomated safety check: PassNo licence6 mo ago
86

Query live GPU inventory, submit an authenticated Itô fixed-rate RFQ, inspect RFQ or procurement status, revoke device credentials, and run explicitly gated node qualification through the separately…

affaan-m/ECC275k1 repo~1.7kAutomated safety check: PassMIT3 days ago
87

Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming).

yzlnew/infra-skills149—~2.4kAutomated safety check: PassNo licence3 mo ago
88

Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs.

K-Dense-AI/scientific-agent-skills48k1 repo~4.5kAutomated safety check: NotesApache-2.02 days ago
89

GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

K-Dense-AI/scientific-agent-skills48k1 repo~3.4kAutomated safety check: PassMIT2 days ago
90

Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.

sickn33/agentic-awesome-skills47k2 repos~3.2kAutomated safety check: PassMITtoday
91

A skill your agent uses when the user asks "why is this B200/Blackwell kernel slow, what should I optimize first, should I fuse or move to Tensor Cores/TMA".

mirage-project/mirage2.5k—~2.2kAutomated safety check: PassApache-2.0today
92

A skill your agent uses when a B200/Blackwell kernel shows wrong results, uncoalesced global memory access, SMEM bank conflicts, a TMA swizzle that mismatches the Tensor Core read, or confused…

mirage-project/mirage2.5k—~1.8kAutomated safety check: PassApache-2.0today
93

A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer…

mirage-project/mirage2.5k—~1.3kAutomated safety check: PassApache-2.0today
94
94.Hyperpod NcclOfficial

Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…

awslabs/agent-plugins915—~3.4kAutomated safety check: PassApache-2.02 days ago
95
95.Local AI AgentsOfficial

Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models.

microsoft/ai-agents-for-beginners77k—~1.3kAutomated safety check: PassMIT19 days ago
96
96.Local AI AgentsOfficial

Build local-first AI agents wey dey run fully for developer workstation wit Microsoft Foundry Local and Qwen function-calling models.

microsoft/ai-agents-for-beginners77k—~1.4kAutomated safety check: PassMIT19 days ago