Search
GPU and accelerator computing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits. | huggingface/ | 11k | 2 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 50 | Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran. | BBuf/ | 938 | — | ~2k | Automated safety check: Pass | No licence | 4 days ago |
| 51 | A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances. | aws/ | 103 | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | today |
| 52 | Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput. | awslabs/ | 916 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | today |
| 53 | 53.Triton Skill Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 54 | Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. | lucifer1004/ | 128 | — | ~1.3k | Automated safety check: Pass | MIT | 5 days ago |
| 55 | 55.Cuda Draft, debug, and measure CUDA kernels and host launch workflows. | sablin39/ | 145 | — | ~990 | Automated safety check: Pass | No licence | 24 days ago |
| 56 | Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably. | Krusty84/ | 106 | — | ~564 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 57 | Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing. | microsoft/ | 22k | — | ~6.5k | Automated safety check: Pass | MIT | today |
| 58 | Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. | intel/ | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 59 | Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model. | Orchestra-Research/ | 13k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 60 | Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery. | Orchestra-Research/ | 13k | 2 repos | ~2.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 61 | Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 62 | Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use. | davila7/ | 33k | 10 repos | ~2.4k | Automated safety check: Pass | MIT | today |
| 63 | Organizes PyTorch training code into LightningModules, DataModules and Trainers, with multi-GPU strategies, callbacks and logging configured. | davila7/ | 33k | 11 repos | ~1.7k | Automated safety check: Pass | MIT | today |
| 64 | Review 3DGS implementation code for correctness, performance bugs, and best practices. | jaccen/ | 161 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 65 | Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering. | Krusty84/ | 106 | — | ~597 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 66 | Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect. | BBuf/ | 938 | — | ~3.9k | Automated safety check: Pass | No licence | 4 days ago |
| 67 | Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons. | Orchestra-Research/ | 13k | 2 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 68 | Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download. | NVIDIA/ | 3.6k | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | today |
| 69 | PyTorch training reference: architecture choice by data type, scaling rules, a training loop, optimizer and learning-rate choices, and fixes for loss spikes or OOM. | Orchestra-Research/ | 13k | 1 repo | ~2.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 70 | Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives. | Orchestra-Research/ | 13k | 4 repos | ~3k | Automated safety check: Warn | MIT | 3 mo ago |
| 71 | Write optimized Triton GPU kernels for deep learning operations. | vipshop/ | 1.3k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 72 | Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA. | BBuf/ | 938 | — | ~7.5k | Automated safety check: Pass | No licence | 4 days ago |
| 73 | Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~709 | Automated safety check: Pass | MIT | today |
| 74 | Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs. | Krusty84/ | 106 | — | ~662 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 75 | Guided step-by-step wizard for producing a new FlyDSL GPU kernel from a requirement: classify the kernel type, pick a skeleton, fill in compute, add control flow / sync / LDS, then test on GPU. | ROCm/ | 291 | — | ~4.6k | Automated safety check: Notes | Unknown | today |
| 76 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 77 | Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX). | facebookexperimental/ | 201 | — | ~644 | Automated safety check: Pass | MIT | today |
| 78 | Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation. | Krusty84/ | 106 | — | ~753 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 79 | Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling. | Orchestra-Research/ | 13k | — | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 80 | Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups. | Orchestra-Research/ | 13k | — | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 81 | Fine-tunes and serves Physical Intelligence's pi0, pi0-fast and pi0.5 robot policies with JAX or PyTorch, including checkpoint conversion and policy servers. | Orchestra-Research/ | 13k | — | ~3.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 82 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups | sgl-project/ | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | today |
| 83 | Estimate peak GPU memory for a DualPipeV training run. An agent skill from mlc-ai/pith-train. | mlc-ai/ | 355 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 84 | 3DGS compression-to-deployment pipeline: quantization (scalar/VQ/mixed-precision), pruning (coreset/adaptive/variational/merge/Bayesian), progressive streaming & LoD, Web/WebGPU/mobile deployment… | jaccen/ | 161 | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 85 | Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs. | Krusty84/ | 106 | — | ~636 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 86 | Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~1.1k | Automated safety check: Pass | No licence | 6 mo ago |
| 87 | 87.Ito Compute Query live GPU inventory, submit an authenticated Itô fixed-rate RFQ, inspect RFQ or procurement status, revoke device credentials, and run explicitly gated node qualification through the separately… | affaan-m/ | 276k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | today |
| 88 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | today |
| 89 | Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming). | yzlnew/ | 149 | — | ~2.4k | Automated safety check: Pass | No licence | 3 mo ago |
| 90 | 90.Modal Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs. | K-Dense-AI/ | 48k | 1 repo | ~4.5k | Automated safety check: Notes | Apache-2.0 | 4 days ago |
| 91 | GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster. | K-Dense-AI/ | 48k | 1 repo | ~3.4k | Automated safety check: Pass | MIT | 4 days ago |
| 92 | Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls. | sickn33/ | 47k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | yesterday |
| 93 | A skill your agent uses when the user asks "why is this B200/Blackwell kernel slow, what should I optimize first, should I fuse or move to Tensor Cores/TMA". | mirage-project/ | 2.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 94 | A skill your agent uses when a B200/Blackwell kernel shows wrong results, uncoalesced global memory access, SMEM bank conflicts, a TMA swizzle that mismatches the Tensor Core read, or confused… | mirage-project/ | 2.5k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 95 | A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer… | mirage-project/ | 2.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 96 | Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures… | awslabs/ | 916 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |