Language
CUDA agent skills, page 4
CUDA skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | 145.Research A skill your agent uses for web searches and current information lookups, finding sources, fact-checking, researching questions, comparing sources, or summarising web pages. | zhongkaifu/ | 557 | — | ~2.3k | Automated safety check: Warn | BSD-3-Clause | today |
| 146 | Verify a CV-CUDA optimization campaign's deterministic definition-of-done and concise versioned MR summary per .agents/guidance/OPTIMIZATIONGUIDELINES.md. | CVCUDA/ | 2.7k | — | ~424 | Automated safety check: Pass | Unknown | 21 days ago |
| 147 | 147.Llama Cpp Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. | Orchestra-Research/ | 13k | 4 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 148 | 148.Setup Guide A skill your agent uses when the user wants to set up LLM training for the first time, or when traininghub is not yet installed/configured in the current environment. | Red-Hat-AI-Innovation-Team/ | 100 | — | ~959 | Automated safety check: Pass | Apache-2.0 | today |
| 149 | 149.Add Tts Model Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. | vllm-project/ | 7.1k | — | ~8.7k | Automated safety check: Pass | Apache-2.0 | today |
| 150 | 150.Model Deploy Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling). | sohu-mptc/ | 107 | — | ~974 | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 151 | 151.Tilelang Skill Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source. | slowlyC/ | 169 | — | ~1.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 152 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM… | vipshop/ | 1.3k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 153 | 153.Add New Model Guided workflow for adding a new model architecture to llama.cpp. | JakeATX/ | 148 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 154 | 154.Mindquantum Build, simulate, and analyze quantum circuits with MindQuantum. | mindspore-ai/ | 101 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 155 | Add, fix, or validate Triton Runner support for an exact Triton version. | toyaix/ | 100 | — | ~1.1k | Automated safety check: Pass | MIT | 21 days ago |
| 156 | 156.Veomni Profile A skill your agent uses for performance profiling and optimization. | ByteDance-Seed/ | 2.2k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 157 | 157.Cmake CMake build options, custom functions, and backend patterns for LuisaCompute. | LuisaGroup/ | 1.1k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 158 | Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed… | pytorch/ | 113 | — | ~2k | Automated safety check: Pass | Unknown | today |
| 159 | 159.Add Sgl Kernel Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks) | sgl-project/ | 37k | 2 repos | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |
| 160 | 160.Debug Cuda Crash Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging | sgl-project/ | 37k | 2 repos | ~4.9k | Automated safety check: Pass | Apache-2.0 | today |
| 161 | Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran. | BBuf/ | 911 | — | ~2k | Automated safety check: Pass | No licence | 2 days ago |
| 162 | Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves… | Red-Hat-AI-Innovation-Team/ | 100 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 163 | 163.Triton Skill Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 164 | 164.Backend AI Guide Expert guide for Backend.AI distributed computing platform. An agent skill from lablup/backend.ai-webui. | lablup/ | 133 | 1 repo | ~1.8k | Automated safety check: Pass | LGPL-3.0 | today |
| 165 | 165.Make Op Add a new CV-CUDA operator end-to-end per .agents/guidance/MAKEOPGUIDELINES.md, with a deterministically-enforced definition-of-done. | CVCUDA/ | 2.7k | — | ~831 | Automated safety check: Pass | Unknown | 21 days ago |
| 166 | A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing… | vipshop/ | 1.3k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 167 | Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. | lucifer1004/ | 127 | — | ~1.3k | Automated safety check: Pass | MIT | 3 days ago |
| 168 | 168.Cuda Draft, debug, and measure CUDA kernels and host launch workflows. | sablin39/ | 145 | — | ~990 | Automated safety check: Pass | No licence | 22 days ago |
| 169 | 169.Extending Ocannl Touch-lists for common OCANNL extension tasks: adding a primitive operation, adding or extending a backend, extending shape inference, and diagnosing output differences between backends. | ahrefs/ | 118 | — | ~728 | Automated safety check: Pass | BSD-2-Clause | yesterday |
| 170 | 170.Init GPU Server Initialize a Draw Things GPU server with GPUScript, including script sync, Docker/CUDA/NVIDIA runtime setup, 7T data disk mounting, mergerfs, and end-to-end GPU verification. | drawthingsai/ | 580 | — | ~2.2k | Automated safety check: Pass | GPL-3.0 | yesterday |
| 171 | 171.Test Change Build and run the unit tests covering changed Celeritas source files. | celeritas-project/ | 105 | — | ~640 | Automated safety check: Pass | Unknown | today |
| 172 | Guidelines for NVIDIA GPU kernel / Triton / Gluon / TileLang / CUDA backend performance work in the FLA repo. | fla-org/ | 5.8k | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 173 | Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons. | Orchestra-Research/ | 13k | 3 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 174 | Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works. | vllm-project/ | 7.1k | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 175 | Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific… | alibaba/ | 161 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 176 | 176.Code Review Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. | JakeATX/ | 148 | — | ~5.2k | Automated safety check: Pass | MIT | yesterday |
| 177 | Review 3DGS implementation code for correctness, performance bugs, and best practices. | jaccen/ | 161 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 178 | 178.Veomni Uv Update A skill your agent uses when updating dependencies managed by uv: bumping a package version, upgrading the uv tool itself, updating torch/CUDA stack, switching transformers version, or regenerating… | ByteDance-Seed/ | 2.2k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 179 | 179.Compiling How to compile Gkeyll libraries, executables, and specific unit or regression test targets on CPU or CUDA. | gkeyllorg/ | 112 | — | ~759 | Automated safety check: Pass | MIT | today |
| 180 | 180.Cuda Skill Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. | slowlyC/ | 169 | — | ~1.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 181 | Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives. | Orchestra-Research/ | 13k | 5 repos | ~3k | Automated safety check: Warn | MIT | 3 mo ago |
| 182 | 182.Debug Debug crashes and test failures via stack-traces, host/device logging, and DSL buffer inspection. | LuisaGroup/ | 1.1k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 183 | 183.AI Review 使用 PaddlePaddle 仓库规则评审 Pull Request 和全仓库代码变更,覆盖正确性、兼容性、算子、分布式、数值、性能、安全、测试、构建和 PR 信息。当需要审查 Paddle 的代码、测试、算子 YAML、C++/CUDA/XPU kernel、Python API、分布式逻辑或 CI 配置时使用。 | PaddlePaddle/ | 24k | — | ~303 | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 184 | 184.Make Op Scaffold Scaffold a new CV-CUDA operator — a complete, wired, building skeleton — and delegate the implementation to a human or another AI. | CVCUDA/ | 2.7k | — | ~306 | Automated safety check: Pass | Unknown | 21 days ago |
| 185 | Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API. | vllm-project/ | 103 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 186 | Workflow for upgrading/integrating cache-dit in SGLang diffusion (multimodalgen): DBCache, DMD calibrator, TaylorSeer, SVDQuant DQ; porting upstream PRs and resolving conflicts against the… | sgl-project/ | 37k | — | ~7.1k | Automated safety check: Pass | Apache-2.0 | today |
| 187 | 187.Safe Debug Rigor Debug / Rigor Audit skill for deep learning research work. | lllllllama/ | 497 | 1 repo | ~522 | Automated safety check: Pass | MIT | 14 days ago |
| 188 | Add or debug an AReno model family, including config conversion, module construction, checkpoint load/save, text or multimodal inference, training backward, tensor parallelism, CUDA graph decode… | inclusionAI/ | 323 | — | ~722 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 189 | 189.Make Op Verify Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate). | CVCUDA/ | 2.7k | — | ~433 | Automated safety check: Pass | Unknown | 21 days ago |
| 190 | Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning. | amd/ | 181 | — | ~1.4k | Automated safety check: Pass | MIT | 10 days ago |
| 191 | Measure TileLang kernel and workflow performance, choose benchmark or timeline tools, and compare implementations fairly. | sablin39/ | 145 | — | ~1.6k | Automated safety check: Pass | No licence | 22 days ago |
| 192 | 192.Debug Test Update .vscode/launch.json to debug a specific CTest test by name. | celeritas-project/ | 105 | — | ~199 | Automated safety check: Pass | Unknown | today |