Search
PyTorch · GPU and accelerator computing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 2 | Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel. | linkedin/ | 6.7k | — | ~799 | Automated safety check: Pass | BSD-2-Clause | 2 days ago |
| 3 | Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. | open-infra-skills/ | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 4 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 5 | Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~4.9k | Automated safety check: Pass | Unknown | today |
| 6 | Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling. | Orchestra-Research/ | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 7 | Debug PyTorch 2 compiler stack failures including Dynamo graph breaks, Inductor codegen errors, AOTAutograd crashes, and accuracy mismatches. | pytorch/ | 104k | — | ~3.5k | Automated safety check: Pass | Unknown | today |
| 8 | Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups. | Orchestra-Research/ | 13k | 5 repos | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 9 | Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. | lucifer1004/ | 128 | — | ~1.3k | Automated safety check: Pass | MIT | 6 days ago |
| 10 | Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. | intel/ | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery. | Orchestra-Research/ | 13k | 2 repos | ~2.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 12 | Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 13 | Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use. | davila7/ | 33k | 10 repos | ~2.4k | Automated safety check: Pass | MIT | today |
| 14 | Organizes PyTorch training code into LightningModules, DataModules and Trainers, with multi-GPU strategies, callbacks and logging configured. | davila7/ | 33k | 11 repos | ~1.7k | Automated safety check: Pass | MIT | today |
| 15 | Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons. | Orchestra-Research/ | 13k | 2 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 16 | PyTorch training reference: architecture choice by data type, scaling rules, a training loop, optimizer and learning-rate choices, and fixes for loss spikes or OOM. | Orchestra-Research/ | 13k | 1 repo | ~2.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 17 | Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives. | Orchestra-Research/ | 13k | 4 repos | ~3k | Automated safety check: Warn | MIT | 3 mo ago |
| 18 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 19 | Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups. | Orchestra-Research/ | 13k | — | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 20 | Fine-tunes and serves Physical Intelligence's pi0, pi0-fast and pi0.5 robot policies with JAX or PyTorch, including checkpoint conversion and policy servers. | Orchestra-Research/ | 13k | — | ~3.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 21 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 22 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.6k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 23 | Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM… | NVIDIA/ | 3.6k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 24 | Run distributed GPU training jobs on CoreWeave with multi-node PyTorch. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 25 | GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. | Mathews-Tom/ | 329 | — | ~3.5k | Automated safety check: Notes | MIT | 4 days ago |
| 26 | 26.Triton Lang Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 252 | — | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |