Topic · AI & LLM Engineering
Best GPU and accelerator computing skills, page 2
GPU and accelerator computing skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count. | Orchestra-Research/ | 13k | 3 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 50 | Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits. | huggingface/ | 11k | 2 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 51 | Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran. | BBuf/ | 911 | — | ~2k | Automated safety check: Pass | No licence | 2 days ago |
| 52 | Organizes PyTorch training code into LightningModules, DataModules and Trainers, with multi-GPU strategies, callbacks and logging configured. | davila7/ | 32k | 14 repos | ~1.7k | Automated safety check: Pass | MIT | today |
| 53 | 53.Triton Skill Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 54 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 55 | Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. | lucifer1004/ | 127 | — | ~1.3k | Automated safety check: Pass | MIT | 3 days ago |
| 56 | 56.Cuda Draft, debug, and measure CUDA kernels and host launch workflows. | sablin39/ | 145 | — | ~990 | Automated safety check: Pass | No licence | 22 days ago |
| 57 | Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably. | Krusty84/ | 106 | — | ~564 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 58 | Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing. | microsoft/ | 22k | — | ~6.5k | Automated safety check: Pass | MIT | today |
| 59 | Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. | intel/ | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 60 | PyTorch training reference: architecture choice by data type, scaling rules, a training loop, optimizer and learning-rate choices, and fixes for loss spikes or OOM. | Orchestra-Research/ | 13k | 2 repos | ~2.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 61 | Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use. | davila7/ | 32k | 10 repos | ~2.4k | Automated safety check: Pass | MIT | today |
| 62 | Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons. | Orchestra-Research/ | 13k | 3 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 63 | A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances. | aws/ | 100 | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | today |
| 64 | Review 3DGS implementation code for correctness, performance bugs, and best practices. | jaccen/ | 161 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 65 | Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering. | Krusty84/ | 106 | — | ~597 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 66 | Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives. | Orchestra-Research/ | 13k | 5 repos | ~3k | Automated safety check: Warn | MIT | 3 mo ago |
| 67 | Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect. | BBuf/ | 911 | — | ~3.9k | Automated safety check: Pass | No licence | 2 days ago |
| 68 | Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download. | NVIDIA/ | 3.5k | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | today |
| 69 | Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups. | Orchestra-Research/ | 13k | 1 repo | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 70 | Fine-tunes and serves Physical Intelligence's pi0, pi0-fast and pi0.5 robot policies with JAX or PyTorch, including checkpoint conversion and policy servers. | Orchestra-Research/ | 13k | 1 repo | ~3.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 71 | Write optimized Triton GPU kernels for deep learning operations. | vipshop/ | 1.3k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 72 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 915 | 1 repo | ~910 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 73 | Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA. | BBuf/ | 911 | — | ~7.5k | Automated safety check: Pass | No licence | 2 days ago |
| 74 | Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~709 | Automated safety check: Pass | MIT | today |
| 75 | Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs. | Krusty84/ | 106 | — | ~662 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 76 | Guided step-by-step wizard for producing a new FlyDSL GPU kernel from a requirement: classify the kernel type, pick a skeleton, fill in compute, add control flow / sync / LDS, then test on GPU. | ROCm/ | 288 | — | ~4.6k | Automated safety check: Notes | Unknown | today |
| 77 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 398 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 78 | Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX). | facebookexperimental/ | 201 | — | ~644 | Automated safety check: Pass | MIT | today |
| 79 | Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation. | Krusty84/ | 106 | — | ~753 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 80 | Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling. | Orchestra-Research/ | 13k | — | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 81 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups | sgl-project/ | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | today |
| 82 | Estimate peak GPU memory for a DualPipeV training run. An agent skill from mlc-ai/pith-train. | mlc-ai/ | 355 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 83 | 3DGS compression-to-deployment pipeline: quantization (scalar/VQ/mixed-precision), pruning (coreset/adaptive/variational/merge/Bayesian), progressive streaming & LoD, Web/WebGPU/mobile deployment… | jaccen/ | 161 | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 84 | Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs. | Krusty84/ | 106 | — | ~636 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 85 | Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~1.1k | Automated safety check: Pass | No licence | 6 mo ago |
| 86 | 86.Ito Compute Query live GPU inventory, submit an authenticated Itô fixed-rate RFQ, inspect RFQ or procurement status, revoke device credentials, and run explicitly gated node qualification through the separately… | affaan-m/ | 275k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | 3 days ago |
| 87 | Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming). | yzlnew/ | 149 | — | ~2.4k | Automated safety check: Pass | No licence | 3 mo ago |
| 88 | 88.Modal Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs. | K-Dense-AI/ | 48k | 1 repo | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 89 | GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster. | K-Dense-AI/ | 48k | 1 repo | ~3.4k | Automated safety check: Pass | MIT | 2 days ago |
| 90 | Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls. | sickn33/ | 47k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | today |
| 91 | A skill your agent uses when the user asks "why is this B200/Blackwell kernel slow, what should I optimize first, should I fuse or move to Tensor Cores/TMA". | mirage-project/ | 2.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 92 | A skill your agent uses when a B200/Blackwell kernel shows wrong results, uncoalesced global memory access, SMEM bank conflicts, a TMA swizzle that mismatches the Tensor Core read, or confused… | mirage-project/ | 2.5k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 93 | A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer… | mirage-project/ | 2.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 94 | Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures… | awslabs/ | 915 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 95 | Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. | microsoft/ | 77k | — | ~1.3k | Automated safety check: Pass | MIT | 19 days ago |
| 96 | Build local-first AI agents wey dey run fully for developer workstation wit Microsoft Foundry Local and Qwen function-calling models. | microsoft/ | 77k | — | ~1.4k | Automated safety check: Pass | MIT | 19 days ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- LLM inference and serving372
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- LLM evaluation308
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23