Language
CUDA agent skills, page 3
CUDA skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 98 | 98.Esmfold2 Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. | JimLiu/ | 227 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 99 | Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness. | stas00/ | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | yesterday |
| 100 | Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI. | greyhaven-ai/ | 1.3k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | today |
| 101 | Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks. | pytorch/ | 5.1k | — | ~2.3k | Automated safety check: Notes | Unknown | today |
| 102 | 102.Refactor Op Find and safely apply per-operator refactoring / redundancy-reduction opportunities in a CV-CUDA operator (near-duplicate Tensor/VarShape kernels, reinvented shared utilities, dead code). | CVCUDA/ | 2.7k | — | ~1.5k | Automated safety check: Pass | Unknown | 21 days ago |
| 103 | 103.Documents Read and write real documents on the device - PDF, XLSX, DOCX, PPTX and CSV. | zhongkaifu/ | 553 | — | ~4.2k | Automated safety check: Pass | BSD-3-Clause | today |
| 104 | Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL… | guqiong96/ | 464 | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 105 | Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. | open-infra-skills/ | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 106 | Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks. | Blackwellboy/ | 135 | — | ~2.1k | Automated safety check: Pass | MIT | yesterday |
| 107 | 107.Graphsignal Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 108 | Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity. | K-Dense-AI/ | 48k | 1 repo | ~3k | Automated safety check: Notes | MIT | 2 days ago |
| 109 | 109.Veomni Debug A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected… | ByteDance-Seed/ | 2.2k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 110 | 110.Cuda CUDA kernel development, debugging, and performance optimization for Claude Code. | technillogue/ | 229 | — | ~2.5k | Automated safety check: Pass | No licence | 9 mo ago |
| 111 | 111.Benchmark Tune A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing… | Mesh-LLM/ | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 112 | Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass. | KernelFlow-ops/ | 212 | — | ~4.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 113 | Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. | wshobson/ | 40k | 1 repo | ~2k | Automated safety check: Pass | MIT | 2 days ago |
| 114 | 114.Cpp Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics. | crazyguitar/ | 290 | — | ~1.8k | Automated safety check: Pass | MIT | 6 days ago |
| 115 | 115.Paddle Debug 在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。 | PaddlePaddle/ | 24k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 116 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 900 | — | ~2.8k | Automated safety check: Pass | No licence | 2 days ago |
| 117 | 117.Optimize Op Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary. | CVCUDA/ | 2.7k | — | ~834 | Automated safety check: Pass | Unknown | 21 days ago |
| 118 | 118.Metal Kernel Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~4.9k | Automated safety check: Pass | Unknown | today |
| 119 | Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered. | meta-pytorch/ | 1.3k | — | ~6.4k | Automated safety check: Pass | BSD-3-Clause | 2 days ago |
| 120 | 120.Chai Structure prediction using Chai-1, a foundation model for molecular structure. | adaptyvbio/ | 163 | 4 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 121 | 121.Ako4all Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup. | TongmingLAIC/ | 369 | — | ~4k | Automated safety check: Pass | MIT | 22 days ago |
| 122 | 122.Cuda Cpp Kernel A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems… | vipshop/ | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 123 | Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator. | inclusionAI/ | 323 | — | ~498 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 124 | 124.Paddle Op Dev PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML… | PaddlePaddle/ | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 125 | A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop… | RightNow-AI/ | 148 | — | ~1.8k | Automated safety check: Pass | MIT | 20 days ago |
| 126 | 126.Review Op Review a CV-CUDA operator end-to-end (support / test / bench / docs coverage). | CVCUDA/ | 2.7k | — | ~481 | Automated safety check: Pass | Unknown | 21 days ago |
| 127 | Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). | matlab/ | 181 | 1 repo | ~3.4k | Automated safety check: Pass | Unknown | 27 days ago |
| 128 | 128.Ncu Report Skill Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill. | mit-han-lab/ | 244 | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 129 | Generate a local phyai environment report for debugging system, Python, CUDA/GPU, dependency, workspace package, git, and PHYAI configuration issues. | mingti-org/ | 128 | — | ~590 | Automated safety check: Pass | MIT | yesterday |
| 130 | Upgrade shared Modal runtime dependencies in kernelbot and verify them end to end. | gpu-mode/ | 114 | — | ~748 | Automated safety check: Pass | Unknown | 19 days ago |
| 131 | Fallback installer for milesdiffusion on a bare CUDA 12.9 Linux GPU box, reproducing the official radixark/milesdiffusion image's package versions and verifying them. | radixark/ | 107 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 132 | 132.Market Data Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares). | zhongkaifu/ | 553 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | today |
| 133 | 133.Fix Env Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any… | evo-design/ | 134 | — | ~2.5k | Automated safety check: Notes | MIT | yesterday |
| 134 | 134.Cutlass Skill Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 135 | 135.GPU Optimization GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI. | spiriMirror/ | 335 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 136 | 136.Teach A Model Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better. | Aseiel/ | 157 | — | ~839 | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 137 | Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands. | brevdev/ | 143 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 138 | 138.Cute Dsl Kernel A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or… | vipshop/ | 1.3k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 139 | 139.Readable Cpp Readable C/C++/Rust/CUDA code rules inspired by The Art of Readable Code. | crazyguitar/ | 290 | — | ~6.4k | Automated safety check: Pass | MIT | 6 days ago |
| 140 | 140.Add Jit Kernel Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module | guqiong96/ | 143 | 1 repo | ~10k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 141 | Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage | maxiaosong1124/ | 129 | — | ~1.6k | Automated safety check: Pass | MIT | 4 mo ago |
| 142 | 142.Tirx Ptx Dialect Register, extend, or audit instructions in the table-driven T.ptx dialect (python/tvm/backend/cuda/ptx/table.py), and move the table to a newer PTX ISA version. | mlc-ai/ | 175 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 143 | 143.New Class Scaffold new Celeritas source and test files with the required copyright header and register them in CMake. | celeritas-project/ | 105 | — | ~496 | Automated safety check: Pass | Unknown | yesterday |
| 144 | 144.Research A skill your agent uses for web searches and current information lookups, finding sources, fact-checking, researching questions, comparing sources, or summarising web pages. | zhongkaifu/ | 553 | — | ~2.3k | Automated safety check: Warn | BSD-3-Clause | today |