Search
Python · GPU and accelerator computing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate… | PaddlePaddle/ | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 2 | Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub. | huggingface/ | 11k | 1 repo | ~7.5k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 3 | Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling. | Orchestra-Research/ | 13k | 7 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 4 | Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback. | microsoft/ | 22k | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 5 | Quick-reference help for fine-tuning language models with Axolotl, covering YAML configs, FSDP, context parallelism, compressed saves and dataset formats. | Orchestra-Research/ | 13k | 9 repos | ~1.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 5 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 7 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module | guqiong96/ | 143 | 1 repo | ~10k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 8 | Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups. | Orchestra-Research/ | 13k | 6 repos | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 9 | 9.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | today |
| 10 | Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model. | Orchestra-Research/ | 13k | 3 repos | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 11 | Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery. | Orchestra-Research/ | 13k | 3 repos | ~2.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 12 | Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits. | huggingface/ | 11k | 2 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 13 | 13.Triton Skill Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 14 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 15 | Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use. | davila7/ | 32k | 10 repos | ~2.4k | Automated safety check: Pass | MIT | today |
| 16 | Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups. | Orchestra-Research/ | 13k | 1 repo | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 17 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 915 | 1 repo | ~910 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 18 | Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA. | BBuf/ | 911 | — | ~7.5k | Automated safety check: Pass | No licence | 2 days ago |
| 19 | Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~709 | Automated safety check: Pass | MIT | today |
| 20 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 398 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 21 | Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX). | facebookexperimental/ | 201 | — | ~644 | Automated safety check: Pass | MIT | today |
| 22 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups | sgl-project/ | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | today |
| 23 | Estimate peak GPU memory for a DualPipeV training run. An agent skill from mlc-ai/pith-train. | mlc-ai/ | 355 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 24 | Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs. | Krusty84/ | 106 | — | ~636 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 25 | Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming). | yzlnew/ | 149 | — | ~2.4k | Automated safety check: Pass | No licence | 3 mo ago |
| 26 | 26.Modal Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs. | K-Dense-AI/ | 48k | 1 repo | ~4.5k | Automated safety check: Notes | Apache-2.0 | 3 days ago |
| 27 | GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster. | K-Dense-AI/ | 48k | 1 repo | ~3.4k | Automated safety check: Pass | MIT | 3 days ago |
| 28 | A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame… | NVIDIA/ | 3.5k | 1 repo | ~2.4k | Automated safety check: Notes | Apache-2.0 | today |
| 29 | Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. | NVIDIA/ | 3.5k | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | today |
| 30 | 30.Runpod A skill your agent uses when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless… | ericrisco/ | 167 | — | ~2.8k | Automated safety check: Pass | MIT | today |
| 31 | 31.Modal Cloud computing platform for running Python on GPUs and serverless infrastructure. | BioTender-max/ | 197 | — | ~3.1k | Automated safety check: Notes | Apache-2.0 | 3 mo ago |
| 32 | GPU-accelerate Python code using CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, and RAFT. | majiayu000/ | 666 | 1 repo | ~8.5k | Automated safety check: Pass | MIT | today |
| 33 | 33.Triton Lang Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |