Search
CUDA · For developers
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate). | CVCUDA/ | 2.7k | — | ~433 | Automated safety check: Pass | Unknown | 24 days ago |
| 98 | Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~709 | Automated safety check: Pass | MIT | today |
| 99 | Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning. | amd/ | 182 | — | ~1.4k | Automated safety check: Pass | MIT | 13 days ago |
| 100 | Measure TileLang kernel and workflow performance, choose benchmark or timeline tools, and compare implementations fairly. | sablin39/ | 145 | — | ~1.6k | Automated safety check: Pass | No licence | 26 days ago |
| 101 | 101.Debug Test Update .vscode/launch.json to debug a specific CTest test by name. | celeritas-project/ | 105 | — | ~199 | Automated safety check: Pass | Unknown | today |
| 102 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 103 | Review PyTorch upstream unit-test (UT) PRs that enable Intel GPU (XPU) on existing tests. | intel/ | 115 | — | ~917 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 104 | 104.Mpk Internals Reference guide for the MPK compilation-to-runtime pipeline. | mirage-project/ | 2.5k | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 105 | 105.Add Model 给 FlashRec 引擎接入一个新模型架构(新的 HF checkpoint / 非 Qwen3 结构)。涵盖模型定义、权重合并加载、FP8 双路径、融合 kernel 接线、CUDA graph 兼容、精度校验、以及压测+trace 验证闭环。当用户要"增加/支持/接入新模型"时使用。 | sohu-mptc/ | 107 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 106 | 106.Add Jit Kernel Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups | sgl-project/ | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | today |
| 107 | 107.App Opinionated app components building on top of ./ui primitives | JakeATX/ | 166 | — | ~146 | Automated safety check: Pass | MIT | yesterday |
| 108 | Optimize MATLAB design files for GPU Coder to generate faster CUDA code. | matlab/ | 1.1k | — | ~4.7k | Automated safety check: Pass | Unknown | 2 days ago |
| 109 | Start, validate, debug, and stop an AReno OpenAI-compatible serving endpoint. | inclusionAI/ | 323 | — | ~409 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 110 | Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server. | vllm-project/ | 102 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | 6 mo ago |
| 111 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 112 | Review a CV-CUDA operator's input-type, layout, dtype, and channel support matrix. | CVCUDA/ | 2.7k | — | ~248 | Automated safety check: Pass | Unknown | 24 days ago |
| 113 | Overview of the main directories and important files in the repository. | spiriMirror/ | 336 | — | ~823 | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 114 | 114.Skippy Prompt A skill your agent uses when running or debugging interactive Skippy prompts against staged serving, including lab sync, native builds, stage startup, the HTTP prompt REPL, and process lifecycle. | Mesh-LLM/ | 3.5k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 115 | A skill your agent uses when working with Slang shaders, shader modules, HLSL-compatible GPU code, graphics pipelines, compute shaders, tessellation, ray tracing, parameter blocks, generics… | github/ | 40k | 1 repo | ~1.8k | Automated safety check: Pass | MIT | 2 days ago |
| 116 | Update tools/scripts/generatebinarybuildmatrix.py when a PyTorch release goes live. | pytorch/ | 113 | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 117 | Review a CV-CUDA operator's test coverage, including C++ correctness, required cross-layout parity, correctness rigor, and the Python API surface. | CVCUDA/ | 2.7k | — | ~264 | Automated safety check: Pass | Unknown | 24 days ago |
| 118 | 118.Vllm Server Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | 2 days ago |
| 119 | A skill your agent uses when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin… | mirage-project/ | 2.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 120 | A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer… | mirage-project/ | 2.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 121 | Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures… | awslabs/ | 916 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 122 | AI demos and GPU compute with Gradio Spaces and Hugging Face Spaces ZeroGPU. | sickn33/ | 47k | 1 repo | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 123 | Install Triton + SageAttention to accelerate ComfyUI (the sageattn attentionmode and inductor torch.compile used by WanVideoWrapper / many video graphs). | artokun/ | 803 | — | ~5k | Automated safety check: Pass | MIT | 6 days ago |
| 124 | Apply Gkeyll naming conventions when creating, editing, or reviewing C, CUDA, or Lua files and code elements. | gkeyllorg/ | 113 | — | ~301 | Automated safety check: Pass | MIT | yesterday |
| 125 | Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 126 | Write, review, and run high-level xTBloom Python GFN2-xTB inference with Calculator, Structure, and BatchCalculator, including single systems, repeated geometry updates, heterogeneous ragged… | jinzhezenggroup/ | 148 | — | ~1.3k | Automated safety check: Pass | LGPL-3.0 | 2 days ago |
| 127 | Install or verify the correct ONNX Runtime build (and the onnx package) for a user's accelerator backend before Quark's ONNX-to-ONNX flow. | amd/ | 182 | — | ~3.2k | Automated safety check: Pass | MIT | 13 days ago |
| 128 | Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data. | NVIDIA/ | 3.6k | 1 repo | ~2.3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 129 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.6k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 130 | Continue KERMT pretraining on a custom SMILES corpus with a groverbase, cmim, or hybrid checkpoint. | NVIDIA/ | 3.6k | 1 repo | ~4.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 131 | Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint. | NVIDIA/ | 3.6k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 132 | Finetune a pretrained KERMT encoder on a labeled CSV. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | 1 repo | ~4.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 133 | Check whether a CV-CUDA operator is READY to optimize (correctness + bench coverage + captured baseline + profiling) per .agents/guidance/OPTIMIZATIONGUIDELINES.md. | CVCUDA/ | 2.7k | — | ~255 | Automated safety check: Pass | Unknown | 24 days ago |
| 134 | 5-stage kernel correctness verification protocol for Triton and CUDA kernels. | ZJLi2013/ | 102 | — | ~702 | Automated safety check: Pass | No licence | 6 mo ago |
| 135 | 135.Xmake XMake build configuration, options, commands, and patterns for LuisaCompute. | LuisaGroup/ | 1.1k | — | ~13k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 136 | 136.Heartmula Set up and run HeartMuLa, the open-source music generation model family (Suno-like). | RedWoodOG/ | 177 | 2 repos | ~1.6k | Automated safety check: Pass | No licence | 4 mo ago |
| 137 | Evaluate whether an existing hot path is a credible NVIDIA Warp candidate. | NVIDIA/ | 3.6k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 138 | Pretrain a fresh KERMT model from scratch on a user-provided corpus. | NVIDIA/ | 3.6k | 1 repo | ~2.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 139 | Bootstrap the KERMT agent environment — verify host docker + nvidia-container-toolkit, build the kermt:latest image from the repo's Dockerfile if it doesn't yet exist, and run a GPU smoke test… | NVIDIA/ | 3.6k | 1 repo | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 140 | 140.Heartmula HeartMuLa: Suno-like song generation from lyrics + tags. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 2 repos | ~1.6k | Automated safety check: Pass | MIT | 3 days ago |
| 141 | Inspect a target ONNX model and prepare metadata for Quark ONNX PTQ planning. | amd/ | 182 | — | ~4.3k | Automated safety check: Pass | MIT | 13 days ago |
| 142 | Host setup for TAO GPU backends. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~3.4k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 143 | A skill your agent uses when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the… | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 144 | A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame… | NVIDIA/ | 3.6k | 1 repo | ~2.4k | Automated safety check: Notes | Apache-2.0 | 2 days ago |