Search
CUDA
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | 145.Oob Detection Detect out-of-bounds memory accesses in CPU or GPU code using static interval analysis and runtime assertions/printfs. | ROCm/ | 291 | — | ~1k | Automated safety check: Notes | Unknown | today |
| 146 | 146.Skippy Prompt A skill your agent uses when running or debugging interactive Skippy prompts against staged serving, including lab sync, native builds, stage startup, the HTTP prompt REPL, and process lifecycle. | Mesh-LLM/ | 3.5k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 147 | A skill your agent uses when working with Slang shaders, shader modules, HLSL-compatible GPU code, graphics pipelines, compute shaders, tessellation, ray tracing, parameter blocks, generics… | github/ | 40k | 1 repo | ~1.8k | Automated safety check: Pass | MIT | 2 days ago |
| 148 | Update tools/scripts/generatebinarybuildmatrix.py when a PyTorch release goes live. | pytorch/ | 113 | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 149 | Review a CV-CUDA operator's test coverage, including C++ correctness, required cross-layout parity, correctness rigor, and the Python API surface. | CVCUDA/ | 2.7k | — | ~264 | Automated safety check: Pass | Unknown | 24 days ago |
| 150 | How to build and run GPU targets under Buck in fbcode. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~998 | Automated safety check: Pass | MIT | today |
| 151 | 151.Vllm Server Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | 2 days ago |
| 152 | A skill your agent uses when the user needs to choose the tile, dtype, ctagroup::1/2, SMEM operand layout, or TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell, or to implement an… | mirage-project/ | 2.5k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 153 | A skill your agent uses when a B200/TIRx/CUDA warp-specialized kernel fails to compile, deadlocks, hits an illegal memory access, produces wrong results, or is "correct but slow". | mirage-project/ | 2.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 154 | A skill your agent uses when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin… | mirage-project/ | 2.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 155 | A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer… | mirage-project/ | 2.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 156 | 156.Mixed Precision Use FP16/BF16 mixed precision to accelerate training and reduce memory. | aiming-lab/ | 15k | — | ~275 | Automated safety check: Pass | MIT | 1 mo ago |
| 157 | Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures… | awslabs/ | 916 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 158 | AI demos and GPU compute with Gradio Spaces and Hugging Face Spaces ZeroGPU. | sickn33/ | 47k | 1 repo | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 159 | Install Triton + SageAttention to accelerate ComfyUI (the sageattn attentionmode and inductor torch.compile used by WanVideoWrapper / many video graphs). | artokun/ | 803 | — | ~5k | Automated safety check: Pass | MIT | 6 days ago |
| 160 | Apply Gkeyll naming conventions when creating, editing, or reviewing C, CUDA, or Lua files and code elements. | gkeyllorg/ | 113 | — | ~301 | Automated safety check: Pass | MIT | yesterday |
| 161 | Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 162 | A skill your agent uses when creating, migrating, or debugging pixi environments, especially for scientific Python, bioinformatics, single-cell analysis, CUDA/PyTorch, Jupyter/VS Code kernels… | xuzhougeng/ | 1k | — | ~3.7k | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 163 | Write, review, and run high-level xTBloom Python GFN2-xTB inference with Calculator, Structure, and BatchCalculator, including single systems, repeated geometry updates, heterogeneous ragged… | jinzhezenggroup/ | 148 | — | ~1.3k | Automated safety check: Pass | LGPL-3.0 | 2 days ago |
| 164 | Install or verify the correct ONNX Runtime build (and the onnx package) for a user's accelerator backend before Quark's ONNX-to-ONNX flow. | amd/ | 182 | — | ~3.2k | Automated safety check: Pass | MIT | 13 days ago |
| 165 | Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data. | NVIDIA/ | 3.6k | 1 repo | ~2.3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 166 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.6k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 167 | Continue KERMT pretraining on a custom SMILES corpus with a groverbase, cmim, or hybrid checkpoint. | NVIDIA/ | 3.6k | 1 repo | ~4.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 168 | Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint. | NVIDIA/ | 3.6k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 169 | Finetune a pretrained KERMT encoder on a labeled CSV. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | 1 repo | ~4.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 170 | A skill your agent uses when writing or debugging nvMolKit Python code for GPU-accelerated RDKit fingerprints, similarity, conformers, clustering, and molecular searches. | NVIDIA/ | 3.6k | 1 repo | ~4.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 171 | Check whether a CV-CUDA operator is READY to optimize (correctness + bench coverage + captured baseline + profiling) per .agents/guidance/OPTIMIZATIONGUIDELINES.md. | CVCUDA/ | 2.7k | — | ~255 | Automated safety check: Pass | Unknown | 24 days ago |
| 172 | 5-stage kernel correctness verification protocol for Triton and CUDA kernels. | ZJLi2013/ | 102 | — | ~702 | Automated safety check: Pass | No licence | 6 mo ago |
| 173 | 173.Xmake XMake build configuration, options, commands, and patterns for LuisaCompute. | LuisaGroup/ | 1.1k | — | ~13k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 174 | Continue pretraining from an existing KERMT checkpoint. An agent skill from NVIDIA-BioNeMo/bionemo-agent-toolkit. | NVIDIA-BioNeMo/ | 479 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 175 | 175.Heartmula Set up and run HeartMuLa, the open-source music generation model family (Suno-like). | RedWoodOG/ | 177 | 2 repos | ~1.6k | Automated safety check: Pass | No licence | 4 mo ago |
| 176 | Evaluate whether an existing hot path is a credible NVIDIA Warp candidate. | NVIDIA/ | 3.6k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 177 | A skill your agent uses when training, evaluating, or exporting Workflow policies with online RSL-RL or RLinf, including RL checkpoint and Workflow handoff. | NVIDIA/ | 3.6k | 1 repo | ~3.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 178 | Run predictions with a finetuned KERMT checkpoint on a SMILES-only CSV. | NVIDIA/ | 3.6k | 1 repo | ~1.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 179 | Pretrain a fresh KERMT model from scratch on a user-provided corpus. | NVIDIA/ | 3.6k | 1 repo | ~2.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 180 | Bootstrap the KERMT agent environment — verify host docker + nvidia-container-toolkit, build the kermt:latest image from the repo's Dockerfile if it doesn't yet exist, and run a GPU smoke test… | NVIDIA/ | 3.6k | 1 repo | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 181 | 181.Heartmula HeartMuLa: Suno-like song generation from lyrics + tags. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 2 repos | ~1.6k | Automated safety check: Pass | MIT | 3 days ago |
| 182 | Inspect a target ONNX model and prepare metadata for Quark ONNX PTQ planning. | amd/ | 182 | — | ~4.3k | Automated safety check: Pass | MIT | 13 days ago |
| 183 | Host setup for TAO GPU backends. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~3.4k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 184 | A skill your agent uses when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the… | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 185 | A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame… | NVIDIA/ | 3.6k | 1 repo | ~2.4k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 186 | DALI imperative dynamic mode (nvidia.dali.experimental.dynamic, ndd): use when working on ndd code or migrating pipelines; skip pipeline-only tasks. | NVIDIA/ | 3.6k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 187 | A skill your agent uses when adding, modifying, optimizing, or debugging CuTile autotuning code. | NVIDIA/ | 3.6k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 188 | Mandatory pre-flight compute resource check before running experiments. | OpenLAIR/ | 1.2k | — | ~1.8k | Automated safety check: Pass | MIT | 23 days ago |
| 189 | A skill your agent uses for CUDA-Q setup, simulation targets, QPU access, and @cudaq.kernel authoring guidance. | NVIDIA/ | 3.6k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 190 | Modify, build, test, debug, and contribute to NVIDIA cuOpt (C++/CUDA, Python, server, CI). | NVIDIA/ | 3.6k | — | ~3.2k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 191 | NVIDIA DeepStream SDK development with Python pyservicemaker API. | NVIDIA/ | 3.6k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 192 | Build deterministic forecast scripts with Earth2Studio (model, data source, IO, inference). | NVIDIA/ | 3.6k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |