Search
CUDA
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang. | sgl-project/ | 37k | 3 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch. | pytorch/ | 104k | 1 repo | ~1.7k | Automated safety check: Pass | Unknown | today |
| 3 | A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes. | PaddlePaddle/ | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 | today |
| 4 | A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate… | PaddlePaddle/ | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | today |
| 5 | Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). | sgl-project/ | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm. | ztxz16/ | 5.1k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 7 | Create a CUDA Python pull request from an approved personal or organization-owned fork, including the GitHub CLI GraphQL fallback for renamed organization-owned forks. | NVIDIA/ | 3.4k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 8 | Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild. | microsoft/ | 22k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 9 | Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 10 | Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands. | microsoft/ | 22k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 11 | 11.Esmfold2 Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. | JimLiu/ | 227 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 12 | Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness. | stas00/ | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | 3 days ago |
| 13 | Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI. | greyhaven-ai/ | 1.3k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 14 | Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks. | pytorch/ | 5.1k | — | ~2.3k | Automated safety check: Notes | Unknown | today |
| 15 | 15.Refactor Op Find and safely apply per-operator refactoring / redundancy-reduction opportunities in a CV-CUDA operator (near-duplicate Tensor/VarShape kernels, reinvented shared utilities, dead code). | CVCUDA/ | 2.7k | — | ~1.5k | Automated safety check: Pass | Unknown | 22 days ago |
| 16 | 16.Documents Read and write real documents on the device - PDF, XLSX, DOCX, PPTX and CSV. | zhongkaifu/ | 559 | — | ~4.2k | Automated safety check: Pass | BSD-3-Clause | today |
| 17 | Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component. | microsoft/ | 22k | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 18 | Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL… | guqiong96/ | 465 | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | 17 days ago |
| 19 | Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. | open-infra-skills/ | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 20 | Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks. | Blackwellboy/ | 135 | — | ~2.1k | Automated safety check: Pass | MIT | yesterday |
| 21 | 21.Graphsignal Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 11 days ago |
| 22 | Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity. | K-Dense-AI/ | 48k | 1 repo | ~3k | Automated safety check: Notes | MIT | 4 days ago |
| 23 | 23.Veomni Debug A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected… | ByteDance-Seed/ | 2.2k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 24 | 24.Cuda CUDA kernel development, debugging, and performance optimization for Claude Code. | technillogue/ | 229 | — | ~2.5k | Automated safety check: Pass | No licence | 9 mo ago |
| 25 | A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing… | Mesh-LLM/ | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 26 | Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass. | KernelFlow-ops/ | 213 | — | ~4.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 27 | 27.Cpp Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics. | crazyguitar/ | 290 | — | ~1.8k | Automated safety check: Pass | MIT | 2 days ago |
| 28 | 28.Paddle Debug 在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。 | PaddlePaddle/ | 24k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 29 | Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server. | huggingface/ | 11k | 3 repos | ~945 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 30 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 925 | — | ~2.8k | Automated safety check: Pass | No licence | 4 days ago |
| 31 | 31.Optimize Op Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary. | CVCUDA/ | 2.7k | — | ~834 | Automated safety check: Pass | Unknown | 22 days ago |
| 32 | Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered. | meta-pytorch/ | 1.3k | — | ~6.4k | Automated safety check: Pass | BSD-3-Clause | yesterday |
| 33 | 33.Metal Kernel Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~4.9k | Automated safety check: Pass | Unknown | today |
| 34 | 34.Ako4all Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup. | TongmingLAIC/ | 369 | — | ~4k | Automated safety check: Pass | MIT | 24 days ago |
| 35 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems… | vipshop/ | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 36 | Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback. | microsoft/ | 22k | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 37 | Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator. | inclusionAI/ | 323 | — | ~498 | Automated safety check: Pass | Apache-2.0 | today |
| 38 | PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML… | PaddlePaddle/ | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 39 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 40 | A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop… | RightNow-AI/ | 148 | — | ~1.8k | Automated safety check: Pass | MIT | 21 days ago |
| 41 | Diagnose and fix Cosmos3 environment, installation, and runtime errors. | NVIDIA/ | 559 | — | ~1.3k | Automated safety check: Notes | Unknown | today |
| 42 | 42.Review Op Review a CV-CUDA operator end-to-end (support / test / bench / docs coverage). | CVCUDA/ | 2.7k | — | ~481 | Automated safety check: Pass | Unknown | 22 days ago |
| 43 | Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill. | mit-han-lab/ | 248 | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 44 | Generate a local phyai environment report for debugging system, Python, CUDA/GPU, dependency, workspace package, git, and PHYAI configuration issues. | mingti-org/ | 130 | — | ~590 | Automated safety check: Pass | MIT | 3 days ago |
| 45 | Upgrade shared Modal runtime dependencies in kernelbot and verify them end to end. | gpu-mode/ | 114 | — | ~748 | Automated safety check: Pass | Unknown | 21 days ago |
| 46 | Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK). | intel/ | 1.6k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 47 | Fallback installer for milesdiffusion on a bare CUDA 12.9 Linux GPU box, reproducing the official radixark/milesdiffusion image's package versions and verifying them. | radixark/ | 107 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 48 | 48.Market Data Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares). | zhongkaifu/ | 559 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | today |