Search
CUDA · For developers
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | 49.Cv Deploy 基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。 | LMIXR/ | 188 | — | ~547 | Automated safety check: Pass | No licence | 12 days ago |
| 50 | 50.Readable Cpp Readable C/C++/Rust/CUDA code rules inspired by The Art of Readable Code. | crazyguitar/ | 290 | — | ~6.4k | Automated safety check: Pass | MIT | yesterday |
| 51 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module | guqiong96/ | 144 | 1 repo | ~10k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 52 | Register, extend, or audit instructions in the table-driven T.ptx dialect (python/tvm/backend/cuda/ptx/table.py), and move the table to a newer PTX ISA version. | mlc-ai/ | 175 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 53 | 53.New Class Scaffold new Celeritas source and test files with the required copyright header and register them in CMake. | celeritas-project/ | 105 | — | ~496 | Automated safety check: Pass | Unknown | today |
| 54 | A skill your agent uses when starting a fresh CompileIQ project, hitting a socket timeout, or before running any other compileiq- skill. | NVIDIA/ | 138 | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | 18 days ago |
| 55 | Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. | vllm-project/ | 7.1k | — | ~8.7k | Automated safety check: Pass | Apache-2.0 | today |
| 56 | Guided workflow for adding a new model architecture to llama.cpp. | JakeATX/ | 166 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 57 | Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source. | slowlyC/ | 169 | — | ~1.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 58 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM… | vipshop/ | 1.3k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 59 | A skill your agent uses for performance profiling and optimization. | ByteDance-Seed/ | 2.2k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 60 | 60.Cmake CMake build options, custom functions, and backend patterns for LuisaCompute. | LuisaGroup/ | 1.1k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 61 | 61.Llama Cpp Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. | Orchestra-Research/ | 13k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 62 | Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed… | pytorch/ | 113 | — | ~2k | Automated safety check: Pass | Unknown | today |
| 63 | A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is… | NVIDIA/ | 138 | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | 18 days ago |
| 64 | Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits. | huggingface/ | 11k | 2 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 65 | Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran. | BBuf/ | 938 | — | ~2k | Automated safety check: Pass | No licence | 6 days ago |
| 66 | Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks) | sgl-project/ | 37k | 2 repos | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |
| 67 | Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging | sgl-project/ | 37k | 2 repos | ~4.9k | Automated safety check: Pass | Apache-2.0 | today |
| 68 | 68.Triton Skill Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 69 | Expert guide for Backend.AI distributed computing platform. An agent skill from lablup/backend.ai-webui. | lablup/ | 133 | 1 repo | ~1.8k | Automated safety check: Pass | LGPL-3.0 | today |
| 70 | 70.Make Op Add a new CV-CUDA operator end-to-end per .agents/guidance/MAKEOPGUIDELINES.md, with a deterministically-enforced definition-of-done. | CVCUDA/ | 2.7k | — | ~831 | Automated safety check: Pass | Unknown | 24 days ago |
| 71 | A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing… | vipshop/ | 1.3k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 72 | Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. | lucifer1004/ | 128 | — | ~1.3k | Automated safety check: Pass | MIT | 6 days ago |
| 73 | 73.Cuda Draft, debug, and measure CUDA kernels and host launch workflows. | sablin39/ | 145 | — | ~990 | Automated safety check: Pass | No licence | 26 days ago |
| 74 | Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired… | NVIDIA/ | 560 | — | ~2.7k | Automated safety check: Pass | Unknown | yesterday |
| 75 | Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing. | microsoft/ | 22k | — | ~6.5k | Automated safety check: Pass | MIT | today |
| 76 | Touch-lists for common OCANNL extension tasks: adding a primitive operation, adding or extending a backend, extending shape inference, and diagnosing output differences between backends. | ahrefs/ | 118 | — | ~728 | Automated safety check: Pass | BSD-2-Clause | 4 days ago |
| 77 | Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. | intel/ | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 78 | A skill your agent uses when composing the Search(...) call and calling .start(). | NVIDIA/ | 138 | — | ~2.2k | Automated safety check: Notes | Apache-2.0 | 18 days ago |
| 79 | Initialize a Draw Things GPU server with GPUScript, including script sync, Docker/CUDA/NVIDIA runtime setup, 7T data disk mounting, mergerfs, and end-to-end GPU verification. | drawthingsai/ | 584 | — | ~2.2k | Automated safety check: Pass | GPL-3.0 | yesterday |
| 80 | Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). | matlab/ | 184 | — | ~3.4k | Automated safety check: Pass | Unknown | yesterday |
| 81 | 81.Test Change Build and run the unit tests covering changed Celeritas source files. | celeritas-project/ | 105 | — | ~640 | Automated safety check: Pass | Unknown | today |
| 82 | 82.Code Review Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. | JakeATX/ | 166 | — | ~5.6k | Automated safety check: Pass | MIT | yesterday |
| 83 | Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific… | alibaba/ | 168 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 84 | Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works. | vllm-project/ | 7.1k | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 85 | Review 3DGS implementation code for correctness, performance bugs, and best practices. | jaccen/ | 161 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 86 | A skill your agent uses when updating dependencies managed by uv: bumping a package version, upgrading the uv tool itself, updating torch/CUDA stack, switching transformers version, or regenerating… | ByteDance-Seed/ | 2.2k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 87 | Guide users through Cosmos3 installation, environment setup, checkpoint downloading, and verification. | NVIDIA/ | 560 | — | ~1.1k | Automated safety check: Notes | Unknown | yesterday |
| 88 | 88.Compiling How to compile Gkeyll libraries, executables, and specific unit or regression test targets on CPU or CUDA. | gkeyllorg/ | 113 | — | ~759 | Automated safety check: Pass | MIT | today |
| 89 | 89.Cuda Skill Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. | slowlyC/ | 169 | — | ~1.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 90 | 90.Debug Debug crashes and test failures via stack-traces, host/device logging, and DSL buffer inspection. | LuisaGroup/ | 1.1k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 91 | 91.AI Review 使用 PaddlePaddle 仓库规则评审 Pull Request 和全仓库代码变更,覆盖正确性、兼容性、算子、分布式、数值、性能、安全、测试、构建和 PR 信息。当需要审查 Paddle 的代码、测试、算子 YAML、C++/CUDA/XPU kernel、Python API、分布式逻辑或 CI 配置时使用。 | PaddlePaddle/ | 24k | — | ~303 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 92 | Scaffold a new CV-CUDA operator — a complete, wired, building skeleton — and delegate the implementation to a human or another AI. | CVCUDA/ | 2.7k | — | ~306 | Automated safety check: Pass | Unknown | 24 days ago |
| 93 | Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs. | facebookexperimental/ | 201 | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 94 | Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API. | vllm-project/ | 102 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 95 | Workflow for upgrading/integrating cache-dit in SGLang diffusion (multimodalgen): DBCache, DMD calibrator, TaylorSeer, SVDQuant DQ; porting upstream PRs and resolving conflicts against the… | sgl-project/ | 37k | — | ~7.1k | Automated safety check: Pass | Apache-2.0 | today |
| 96 | 96.Safe Debug Rigor Debug / Rigor Audit skill for deep learning research work. | lllllllama/ | 497 | 1 repo | ~522 | Automated safety check: Pass | MIT | 18 days ago |