Search
GPU and accelerator computing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF. | huggingface/ | 11k | 3 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 2 | Runs an acceptance-style evaluation of an ML GPU cluster: environment dump, matmul FLOPS, all-reduce, disk benchmarks and a dated markdown report. | stas00/ | 19k | — | ~25k | Automated safety check: Notes | CC-BY-SA-4.0 | yesterday |
| 3 | A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate… | PaddlePaddle/ | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 4 | Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). | sgl-project/ | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 5 | Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild. | microsoft/ | 22k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 6 | Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 7 | Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends. | huggingface/ | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 8 | Adds Liger Kernel support for a new HuggingFace Transformers model, or modifies existing monkey-patching. | linkedin/ | 6.6k | — | ~1.3k | Automated safety check: Pass | BSD-2-Clause | today |
| 9 | Operates the Inspire ML platform through its local `inspire` CLI: picking account, workspace and resources, launching notebooks, jobs and services, then cleaning up. | realZillionX/ | 548 | — | ~1.4k | Automated safety check: Pass | MIT | 10 days ago |
| 10 | Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel. | linkedin/ | 6.6k | — | ~799 | Automated safety check: Pass | BSD-2-Clause | today |
| 11 | Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA… | fla-org/ | 5.8k | — | ~4.2k | Automated safety check: Pass | MIT | yesterday |
| 12 | Optimizes the performance of existing Liger Kernel Triton kernels. | linkedin/ | 6.6k | — | ~1.5k | Automated safety check: Pass | BSD-2-Clause | today |
| 13 | Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark. | wshobson/ | 40k | 1 repo | ~2k | Automated safety check: Pass | MIT | 2 days ago |
| 14 | Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths. | BBuf/ | 900 | — | ~2.5k | Automated safety check: Pass | No licence | 2 days ago |
| 15 | Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit. | vllm-project/ | 2.9k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 16 | Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. | open-infra-skills/ | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 17 | 17.Graphsignal Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 18 | 18.Cuda CUDA kernel development, debugging, and performance optimization for Claude Code. | technillogue/ | 229 | — | ~2.5k | Automated safety check: Pass | No licence | 9 mo ago |
| 19 | Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass. | KernelFlow-ops/ | 212 | — | ~4.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 20 | Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. | wshobson/ | 40k | 1 repo | ~2k | Automated safety check: Pass | MIT | 2 days ago |
| 21 | Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno. | inclusionAI/ | 323 | — | ~486 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 22 | Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub. | huggingface/ | 11k | 1 repo | ~7.5k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 23 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 900 | — | ~2.8k | Automated safety check: Pass | No licence | 2 days ago |
| 24 | 24.Metal Kernel Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~4.9k | Automated safety check: Pass | Unknown | today |
| 25 | Guides building on the 0G Compute Network, a decentralized GPU marketplace for AI inference and fine-tuning, with SDK patterns and CLI commands. | internet-court/ | 6.4k | 1 repo | ~1.9k | Automated safety check: Pass | Unknown | 1 mo ago |
| 26 | Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling. | Orchestra-Research/ | 13k | 7 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 27 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems… | vipshop/ | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 28 | Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback. | microsoft/ | 22k | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 29 | A skill your agent uses when adding a new optimized kernel or operator to veomni/ops/. | ByteDance-Seed/ | 2.2k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 30 | Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill. | mit-han-lab/ | 244 | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 31 | Quick-reference help for fine-tuning language models with Axolotl, covering YAML configs, FSDP, context parallelism, compressed saves and dataset formats. | Orchestra-Research/ | 13k | 9 repos | ~1.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 32 | Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks. | BBuf/ | 900 | — | ~4.5k | Automated safety check: Pass | No licence | 2 days ago |
| 33 | GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI. | spiriMirror/ | 335 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 34 | Debug PyTorch 2 compiler stack failures including Dynamo graph breaks, Inductor codegen errors, AOTAutograd crashes, and accuracy mismatches. | pytorch/ | 104k | — | ~3.5k | Automated safety check: Pass | Unknown | today |
| 35 | Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs. | NVIDIA/ | 18k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 36 | Serverless GPU cloud platform for running ML workloads. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 5 repos | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 37 | 37.Nemo Curator GPU-accelerated data curation for LLM training. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 38 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 5 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 39 | Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation. | guqiong96/ | 464 | 1 repo | ~831 | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 40 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module | guqiong96/ | 143 | 1 repo | ~10k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 41 | Shows how to build and run ONNX Runtime WebGPU provider tests on Linux with no GPU, using the Mesa lavapipe software Vulkan adapter, and where that approach falls short. | microsoft/ | 22k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 42 | Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost. | Orchestra-Research/ | 13k | 4 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 43 | Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups. | Orchestra-Research/ | 13k | 6 repos | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 44 | 44.Cv Deploy 基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。 | LMIXR/ | 126 | — | ~547 | Automated safety check: Pass | No licence | 8 days ago |
| 45 | Run jcm on a Kubernetes GPU cluster — generate benchmark and production Job manifests, pick a comparable GPU, survive eviction, collect results. | climate-analytics-lab/ | 108 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |
| 46 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 47 | 47.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | today |
| 48 | Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model. | Orchestra-Research/ | 13k | 3 repos | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |