Language
CUDA agent skills, page 5
CUDA skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 193 | Review a CV-CUDA operator's BENCHMARK coverage — drivers, layout axis, baselines, the basic-tier floor, row counts, and coverage statistics. | CVCUDA/ | 2.7k | — | ~274 | Automated safety check: Pass | Unknown | 21 days ago |
| 194 | Optimize MATLAB design files for GPU Coder to generate faster CUDA code. | matlab/ | 1.1k | — | ~4.7k | Automated safety check: Pass | Unknown | 7 days ago |
| 195 | 195.GPU Backend pyqula's CPU/GPU switch (src/pyqula/gpu.py), how a routine is routed onto the device, per-call precision, and the tiered porting plan in documentation/gpuportingplan.md. | joselado/ | 145 | — | ~679 | Automated safety check: Pass | GPL-3.0 | today |
| 196 | 196.Add Model 给 FlashRec 引擎接入一个新模型架构(新的 HF checkpoint / 非 Qwen3 结构)。涵盖模型定义、权重合并加载、FP8 双路径、融合 kernel 接线、CUDA graph 兼容、精度校验、以及压测+trace 验证闭环。当用户要"增加/支持/接入新模型"时使用。 | sohu-mptc/ | 107 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 197 | 197.Add Jit Kernel Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups | sgl-project/ | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | today |
| 198 | Review a CV-CUDA operator's DOCS & API artifacts — operatorlist row, Python autofunction (fn + into), Limitations-table-vs-code consistency, docstrings, and SPDX headers. | CVCUDA/ | 2.7k | — | ~270 | Automated safety check: Pass | Unknown | 21 days ago |
| 199 | Builds the code for a frozen research experiment test-first, with leakage controls, seed handling and saved evidence so results can be rerun and audited. | Light0305/ | 640 | — | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 200 | 200.App Opinionated app components building on top of ./ui primitives | JakeATX/ | 148 | — | ~146 | Automated safety check: Pass | MIT | yesterday |
| 201 | Start, validate, debug, and stop an AReno OpenAI-compatible serving endpoint. | inclusionAI/ | 323 | — | ~409 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 202 | Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server. | vllm-project/ | 103 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | 6 mo ago |
| 203 | 203.Optimize For GPU GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster. | K-Dense-AI/ | 48k | 1 repo | ~3.4k | Automated safety check: Pass | MIT | 2 days ago |
| 204 | 204.Torchdrug Builds and troubleshoots TorchDrug 0.2.1 workflows for molecular graphs, property prediction, self-supervised pretraining, molecule generation, retrosynthesis, protein representation learning, and… | K-Dense-AI/ | 48k | 1 repo | ~3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 205 | Review a CV-CUDA operator's input-type, layout, dtype, and channel support matrix. | CVCUDA/ | 2.7k | — | ~248 | Automated safety check: Pass | Unknown | 21 days ago |
| 206 | Overview of the main directories and important files in the repository. | spiriMirror/ | 335 | — | ~823 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 207 | 207.Nvmolkit Usage Write code that calls the installed nvMolKit Python API for GPU-accelerated, batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine similarity, ETKDG conformer embedding, MMFF/UFF… | NVIDIA-BioNeMo/ | 478 | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 208 | 208.Oob Detection Detect out-of-bounds memory accesses in CPU or GPU code using static interval analysis and runtime assertions/printfs. | ROCm/ | 287 | — | ~1k | Automated safety check: Notes | Unknown | today |
| 209 | 209.Skippy Prompt A skill your agent uses when running or debugging interactive Skippy prompts against staged serving, including lab sync, native builds, stage startup, the HTTP prompt REPL, and process lifecycle. | Mesh-LLM/ | 3.5k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 210 | Update tools/scripts/generatebinarybuildmatrix.py when a PyTorch release goes live. | pytorch/ | 113 | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 211 | Review a CV-CUDA operator's test coverage, including C++ correctness, required cross-layout parity, correctness rigor, and the Python API surface. | CVCUDA/ | 2.7k | — | ~264 | Automated safety check: Pass | Unknown | 21 days ago |
| 212 | Apply Gkeyll naming conventions when creating, editing, or reviewing C, CUDA, or Lua files and code elements. | gkeyllorg/ | 112 | — | ~301 | Automated safety check: Pass | MIT | today |
| 213 | 213.Vllm Server Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 214 | A skill your agent uses when the user needs to choose the tile, dtype, ctagroup::1/2, SMEM operand layout, or TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell, or to implement an… | mirage-project/ | 2.5k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 215 | A skill your agent uses when a B200/TIRx/CUDA warp-specialized kernel fails to compile, deadlocks, hits an illegal memory access, produces wrong results, or is "correct but slow". | mirage-project/ | 2.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 216 | A skill your agent uses when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin… | mirage-project/ | 2.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 217 | A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer… | mirage-project/ | 2.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 218 | 218.Mixed Precision Use FP16/BF16 mixed precision to accelerate training and reduce memory. | aiming-lab/ | 15k | — | ~275 | Automated safety check: Pass | MIT | 1 mo ago |
| 219 | AI demos and GPU compute with Gradio Spaces and Hugging Face Spaces ZeroGPU. | sickn33/ | 47k | 1 repo | ~4.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 220 | Install Triton + SageAttention to accelerate ComfyUI (the sageattn attentionmode and inductor torch.compile used by WanVideoWrapper / many video graphs). | artokun/ | 793 | — | ~5k | Automated safety check: Pass | MIT | 2 days ago |
| 221 | Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 3 days ago |
| 222 | A skill your agent uses when creating, migrating, or debugging pixi environments, especially for scientific Python, bioinformatics, single-cell analysis, CUDA/PyTorch, Jupyter/VS Code kernels… | xuzhougeng/ | 1k | — | ~3.7k | Automated safety check: Pass | AGPL-3.0 | today |
| 223 | Write, review, and run high-level xTBloom Python GFN2-xTB inference with Calculator, Structure, and BatchCalculator, including single systems, repeated geometry updates, heterogeneous ragged… | jinzhezenggroup/ | 148 | — | ~1.3k | Automated safety check: Pass | LGPL-3.0 | 2 days ago |
| 224 | Install or verify the correct ONNX Runtime build (and the onnx package) for a user's accelerator backend before Quark's ONNX-to-ONNX flow. | amd/ | 181 | — | ~3.2k | Automated safety check: Pass | MIT | 10 days ago |
| 225 | Check whether a CV-CUDA operator is READY to optimize (correctness + bench coverage + captured baseline + profiling) per .agents/guidance/OPTIMIZATIONGUIDELINES.md. | CVCUDA/ | 2.7k | — | ~255 | Automated safety check: Pass | Unknown | 21 days ago |
| 226 | 5-stage kernel correctness verification protocol for Triton and CUDA kernels. | ZJLi2013/ | 102 | — | ~702 | Automated safety check: Pass | No licence | 6 mo ago |
| 227 | 227.Xmake XMake build configuration, options, commands, and patterns for LuisaCompute. | LuisaGroup/ | 1.1k | — | ~13k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 228 | 228.Heartmula HeartMuLa: Suno-like song generation from lyrics + tags. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 125 | 3 repos | ~1.6k | Automated safety check: Pass | MIT | today |
| 229 | Continue pretraining from an existing KERMT checkpoint. An agent skill from NVIDIA-BioNeMo/bionemo-agent-toolkit. | NVIDIA-BioNeMo/ | 478 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 230 | 230.Heartmula Set up and run HeartMuLa, the open-source music generation model family (Suno-like). | RedWoodOG/ | 177 | 2 repos | ~1.6k | Automated safety check: Pass | No licence | 4 mo ago |
| 231 | Inspect a target ONNX model and prepare metadata for Quark ONNX PTQ planning. | amd/ | 181 | — | ~4.3k | Automated safety check: Pass | MIT | 10 days ago |
| 232 | Mandatory pre-flight compute resource check before running experiments. | OpenLAIR/ | 1.2k | — | ~1.8k | Automated safety check: Pass | MIT | 20 days ago |
| 233 | 233.Project Map Maps every HOT-Step CPP feature to its route file, service, UI folder, and engine subsystem, including port topology and the browser-to-engine request path. | scragnog/ | 170 | — | ~5.4k | Automated safety check: Notes | MIT | 2 days ago |
| 234 | Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent. | amd/ | 181 | — | ~4.8k | Automated safety check: Pass | MIT | 10 days ago |
| 235 | Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. | amd/ | 181 | — | ~1.9k | Automated safety check: Notes | MIT | 10 days ago |
| 236 | Install or verify the correct PyTorch build for a user's accelerator backend before Quark installation. | amd/ | 181 | — | ~1.6k | Automated safety check: Pass | MIT | 10 days ago |
| 237 | Diagnoses and resolves MCP server registration failures, GPU detection, BigQuery authentication, index build failures, import errors, search quality issues, and performance problems. | RobThePCGuy/ | 196 | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 238 | Generate C/C++ or CUDA code from an AI model (PyTorch, LiteRT) using MATLAB Coder or GPU Coder. | matlab/ | 1.1k | — | ~2.8k | Automated safety check: Pass | Unknown | 7 days ago |
| 239 | Generate, verify, refine, and accelerate C/C++ or CUDA code from MATLAB with MATLAB Coder, Embedded Coder, GPU Coder, or MATLAB Test. | matlab/ | 1.1k | — | ~4.2k | Automated safety check: Pass | Unknown | 7 days ago |
| 240 | 240.Mat Lammps Md Build and run LAMMPS molecular dynamics with isolated MLIP-specific binaries (MACE, MatGL/CHGNet, FairChem) to avoid Python and Torch stack conflicts. | learningmatter-mit/ | 175 | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |