Search
CUDA
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | A skill your agent uses when composing the Search(...) call and calling .start(). | NVIDIA/ | 138 | — | ~2.2k | Automated safety check: Notes | Apache-2.0 | 17 days ago |
| 98 | Initialize a Draw Things GPU server with GPUScript, including script sync, Docker/CUDA/NVIDIA runtime setup, 7T data disk mounting, mergerfs, and end-to-end GPU verification. | drawthingsai/ | 584 | — | ~2.2k | Automated safety check: Pass | GPL-3.0 | yesterday |
| 99 | Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). | matlab/ | 184 | — | ~3.4k | Automated safety check: Pass | Unknown | yesterday |
| 100 | 100.Test Change Build and run the unit tests covering changed Celeritas source files. | celeritas-project/ | 105 | — | ~640 | Automated safety check: Pass | Unknown | today |
| 101 | 101.Code Review Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. | JakeATX/ | 166 | — | ~5.6k | Automated safety check: Pass | MIT | yesterday |
| 102 | Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific… | alibaba/ | 168 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 103 | Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works. | vllm-project/ | 7.1k | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 104 | Review 3DGS implementation code for correctness, performance bugs, and best practices. | jaccen/ | 161 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 105 | 105.Veomni Uv Update A skill your agent uses when updating dependencies managed by uv: bumping a package version, upgrading the uv tool itself, updating torch/CUDA stack, switching transformers version, or regenerating… | ByteDance-Seed/ | 2.2k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 106 | Guide users through Cosmos3 installation, environment setup, checkpoint downloading, and verification. | NVIDIA/ | 560 | — | ~1.1k | Automated safety check: Notes | Unknown | yesterday |
| 107 | 107.Compiling How to compile Gkeyll libraries, executables, and specific unit or regression test targets on CPU or CUDA. | gkeyllorg/ | 113 | — | ~759 | Automated safety check: Pass | MIT | today |
| 108 | 108.Cuda Skill Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. | slowlyC/ | 169 | — | ~1.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 109 | 109.Debug Debug crashes and test failures via stack-traces, host/device logging, and DSL buffer inspection. | LuisaGroup/ | 1.1k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 110 | 110.AI Review 使用 PaddlePaddle 仓库规则评审 Pull Request 和全仓库代码变更,覆盖正确性、兼容性、算子、分布式、数值、性能、安全、测试、构建和 PR 信息。当需要审查 Paddle 的代码、测试、算子 YAML、C++/CUDA/XPU kernel、Python API、分布式逻辑或 CI 配置时使用。 | PaddlePaddle/ | 24k | — | ~303 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 111 | Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons. | Orchestra-Research/ | 13k | 2 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 112 | 112.Make Op Scaffold Scaffold a new CV-CUDA operator — a complete, wired, building skeleton — and delegate the implementation to a human or another AI. | CVCUDA/ | 2.7k | — | ~306 | Automated safety check: Pass | Unknown | 24 days ago |
| 113 | Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs. | facebookexperimental/ | 201 | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 114 | Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives. | Orchestra-Research/ | 13k | 4 repos | ~3k | Automated safety check: Warn | MIT | 3 mo ago |
| 115 | 115.Cuda Streams Manage tensor lifetimes across CUDA streams and decide whether recordstream is avoidable. | pytorch/ | 104k | — | ~796 | Automated safety check: Pass | Unknown | today |
| 116 | Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API. | vllm-project/ | 102 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 117 | Workflow for upgrading/integrating cache-dit in SGLang diffusion (multimodalgen): DBCache, DMD calibrator, TaylorSeer, SVDQuant DQ; porting upstream PRs and resolving conflicts against the… | sgl-project/ | 37k | — | ~7.1k | Automated safety check: Pass | Apache-2.0 | today |
| 118 | 118.Safe Debug Rigor Debug / Rigor Audit skill for deep learning research work. | lllllllama/ | 497 | 1 repo | ~522 | Automated safety check: Pass | MIT | 17 days ago |
| 119 | Add or debug an AReno model family, including config conversion, module construction, checkpoint load/save, text or multimodal inference, training backward, tensor parallelism, CUDA graph decode… | inclusionAI/ | 323 | — | ~722 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 120 | 120.Make Op Verify Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate). | CVCUDA/ | 2.7k | — | ~433 | Automated safety check: Pass | Unknown | 24 days ago |
| 121 | Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~709 | Automated safety check: Pass | MIT | today |
| 122 | Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning. | amd/ | 182 | — | ~1.4k | Automated safety check: Pass | MIT | 13 days ago |
| 123 | Measure TileLang kernel and workflow performance, choose benchmark or timeline tools, and compare implementations fairly. | sablin39/ | 145 | — | ~1.6k | Automated safety check: Pass | No licence | 26 days ago |
| 124 | 124.Debug Test Update .vscode/launch.json to debug a specific CTest test by name. | celeritas-project/ | 105 | — | ~199 | Automated safety check: Pass | Unknown | today |
| 125 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 126 | Review PyTorch upstream unit-test (UT) PRs that enable Intel GPU (XPU) on existing tests. | intel/ | 115 | — | ~917 | Automated safety check: Pass | Apache-2.0 | today |
| 127 | 127.Mpk Internals Reference guide for the MPK compilation-to-runtime pipeline. | mirage-project/ | 2.5k | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 128 | Review a CV-CUDA operator's BENCHMARK coverage — drivers, layout axis, baselines, the basic-tier floor, row counts, and coverage statistics. | CVCUDA/ | 2.7k | — | ~274 | Automated safety check: Pass | Unknown | 24 days ago |
| 129 | 129.GPU Backend pyqula's CPU/GPU switch (src/pyqula/gpu.py), how a routine is routed onto the device, per-call precision, and the tiered porting plan in documentation/gpuportingplan.md. | joselado/ | 145 | — | ~679 | Automated safety check: Pass | GPL-3.0 | 4 days ago |
| 130 | 130.Add Model 给 FlashRec 引擎接入一个新模型架构(新的 HF checkpoint / 非 Qwen3 结构)。涵盖模型定义、权重合并加载、FP8 双路径、融合 kernel 接线、CUDA graph 兼容、精度校验、以及压测+trace 验证闭环。当用户要"增加/支持/接入新模型"时使用。 | sohu-mptc/ | 107 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 131 | 131.Add Jit Kernel Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups | sgl-project/ | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | today |
| 132 | Review a CV-CUDA operator's DOCS & API artifacts — operatorlist row, Python autofunction (fn + into), Limitations-table-vs-code consistency, docstrings, and SPDX headers. | CVCUDA/ | 2.7k | — | ~270 | Automated safety check: Pass | Unknown | 24 days ago |
| 133 | 133.App Opinionated app components building on top of ./ui primitives | JakeATX/ | 166 | — | ~146 | Automated safety check: Pass | MIT | yesterday |
| 134 | Builds the code for a frozen research experiment test-first, with leakage controls, seed handling and saved evidence so results can be rerun and audited. | Light0305/ | 640 | — | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 135 | Optimize MATLAB design files for GPU Coder to generate faster CUDA code. | matlab/ | 1.1k | — | ~4.7k | Automated safety check: Pass | Unknown | 2 days ago |
| 136 | Start, validate, debug, and stop an AReno OpenAI-compatible serving endpoint. | inclusionAI/ | 323 | — | ~409 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 137 | Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server. | vllm-project/ | 102 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | 6 mo ago |
| 138 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 139 | 139.Optimize For GPU GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster. | K-Dense-AI/ | 48k | 1 repo | ~3.4k | Automated safety check: Pass | MIT | 6 days ago |
| 140 | 140.Torchdrug Builds and troubleshoots TorchDrug 0.2.1 workflows for molecular graphs, property prediction, self-supervised pretraining, molecule generation, retrosynthesis, protein representation learning, and… | K-Dense-AI/ | 48k | 1 repo | ~3k | Automated safety check: Notes | Apache-2.0 | 6 days ago |
| 141 | Turns a parquet of image file paths into a parquet of embeddings with CLIP, SigLIP or a TAO checkpoint, using the TAO Data Services container, ahead of neighbor mining. | NVIDIA/ | 3.6k | — | ~2k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 142 | Review a CV-CUDA operator's input-type, layout, dtype, and channel support matrix. | CVCUDA/ | 2.7k | — | ~248 | Automated safety check: Pass | Unknown | 24 days ago |
| 143 | Overview of the main directories and important files in the repository. | spiriMirror/ | 336 | — | ~823 | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 144 | 144.Nvmolkit Usage Write code that calls the installed nvMolKit Python API for GPU-accelerated, batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine similarity, ETKDG conformer embedding, MMFF/UFF… | NVIDIA-BioNeMo/ | 479 | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |