Language
CUDA agent skills, page 6
CUDA skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 241 | Generate, verify, refine, and accelerate C/C++ or CUDA code from MATLAB with MATLAB Coder, Embedded Coder, GPU Coder, or MATLAB Test. | matlab/ | 1.1k | — | ~4.2k | Automated safety check: Pass | Unknown | yesterday |
| 242 | 242.Mat Lammps Md Build and run LAMMPS molecular dynamics with isolated MLIP-specific binaries (MACE, MatGL/CHGNet, FairChem) to avoid Python and Torch stack conflicts. | learningmatter-mit/ | 176 | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 243 | Generate inorganic material structures using MatterGen, a diffusion-based generative model. | learningmatter-mit/ | 176 | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 244 | 244.Uma Surface Md Run and diagnose governed ASE plus fairchem UMA molecular dynamics on standardized adsorbed slabs, including NVT/NVE surface trajectories, pre-relaxation, periodic-boundary and collision checks… | Tai609/ | 100 | — | ~1.6k | Automated safety check: Pass | Unknown | 1 mo ago |
| 245 | How HOT-Step's custom flash-attention training ops (GGMLOPFLASHATTNTRAIN/BACK) work, what the AS1.5 DiT trainer campaign proved and disproved, and the exact contract for porting flash mode to the… | scragnog/ | 171 | — | ~5.4k | Automated safety check: Pass | MIT | yesterday |
| 246 | Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). | majiayu000/ | 666 | 1 repo | ~4.6k | Automated safety check: Pass | MIT | today |
| 247 | 247.Qmd MCP Skill Use a local QMD knowledge base through UXC over MCP stdio, with daemon-backed session reuse and typed retrieval flows that avoid repeated model warmup and unnecessary query-expansion latency. | holon-run/ | 116 | — | ~1.3k | Automated safety check: Pass | MIT | 23 days ago |
| 248 | 248.Flash Attention Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs. | AlexAI-MCP/ | 135 | — | ~1.2k | Automated safety check: Pass | MIT | 6 mo ago |
| 249 | Diagnose and fix CoreWeave GPU scheduling, pod, and networking errors. | jeremylongshore/ | 2.8k | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 250 | Upgrade CoreWeave deployments and migrate between GPU types. | jeremylongshore/ | 2.8k | — | ~1.9k | Automated safety check: Pass | MIT | today |
| 251 | 251.GPU Optimizer GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. | Mathews-Tom/ | 328 | — | ~3.5k | Automated safety check: Notes | MIT | 2 days ago |
| 252 | 252.Discoart Use this repo skill for DiscoArt image generation, configuration/prompt scheduling, CLI, Jina serving, Docker runtime planning, and troubleshooting. | VectorSpaceLab/ | 328 | — | ~1.3k | Automated safety check: Pass | Unknown | 1 mo ago |
| 253 | 253.Hunyuan Video Use this operating skill for Tencent-Hunyuan/HunyuanVideo text-to-video setup, checkpoint layout, inference commands, Gradio launch, and CUDA/FP8/xDiT troubleshooting. | VectorSpaceLab/ | 328 | — | ~1.1k | Automated safety check: Pass | Unknown | 1 mo ago |
| 254 | 254.Make It 3D Use this repo skill for Make-It-3D single-image 3D creation, including CUDA asset setup, alpha-image validation, coarse NeRF optimization, refinement, rendering, export, and troubleshooting. | VectorSpaceLab/ | 328 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 255 | 255.Mining Backends Use this quip-miner sub-skill for CPU, CUDA, Metal, Modal, and QPU mining commands, backend dependencies, QPU budgets, unified streaming, and PoW/mempool scheduling. | VectorSpaceLab/ | 328 | — | ~753 | Automated safety check: Pass | AGPL-3.0 | 1 mo ago |
| 256 | 256.Nuplan Devkit A skill your agent uses for Motional nuPlan autonomous-driving planning workflows: dataset and map access, scenario filtering, planner implementation, open- or closed-loop simulation, metrics… | VectorSpaceLab/ | 328 | — | ~1.2k | Automated safety check: Pass | Unknown | 1 mo ago |
| 257 | A skill your agent uses for Anomalib benchmark pipelines, tiled ensemble workflows, and advanced pipeline orchestration helpers while keeping experimental execution paths explicit. | VectorSpaceLab/ | 328 | — | ~952 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 258 | 258.Quip Miner Use this repo skill for QuIP quip-miner, the Substrate-integrated quantum mining CLI, when configuring or operating CPU/CUDA/Metal/Modal/QPU miners, managing hybrid wallets/bootstrap/identity… | VectorSpaceLab/ | 328 | — | ~1.4k | Automated safety check: Pass | AGPL-3.0 | 1 mo ago |
| 259 | Use this sub-skill for Torch-TensorRT runtime performance controls, CUDA Graphs, output allocation, caches, TensorRT-RTX runtime settings, mutable modules, refit, weight streaming, and benchmark… | VectorSpaceLab/ | 328 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 260 | 260.Swin Transformer Use this repo skill for Microsoft Swin-Transformer image-classification model, config, data, checkpoint, SimMIM, Swin-MoE, and optional CUDA acceleration workflows. | VectorSpaceLab/ | 328 | — | ~1.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 261 | 261.Torch Tensorrt A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging… | VectorSpaceLab/ | 328 | — | ~1.5k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 262 | Generate novel crystal structures and molecules using ADiT (All-atom Diffusion Transformer), a unified latent diffusion model. | learningmatter-mit/ | 176 | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 263 | 263.Optimize For GPU GPU-accelerate Python code using CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, and RAFT. | majiayu000/ | 666 | 1 repo | ~8.5k | Automated safety check: Pass | MIT | today |
| 264 | AI for Science 场景下的昇腾 NPU Profiling 采集与性能分析 Skill,用于在华为 Ascend NPU 上使用 torchnpu.profiler 采集 L0、L1、L2 级性能数据,分析训练或推理中的算子耗时、调用栈、内存与瓶颈,并指导后续调优。 | ascend-ai-coding/ | 174 | — | ~3k | Automated safety check: Pass | No licence | today |
| 265 | Ankh 蛋白质语言模型昇腾 NPU 迁移 Skill,适用于 Ankh base/large、Ankh3 large/XL 以及同类基于 HuggingFace Transformers 与 PyTorch 的蛋白模型从 CUDA/GPU 到华为 Ascend NPU 的环境检查、代码适配、权重加载、验证脚本补齐与文档沉淀。 | ascend-ai-coding/ | 174 | — | ~2.1k | Automated safety check: Pass | No licence | today |
| 266 | Boltz2 蛋白质结构预测模型的昇腾 NPU 迁移与复现 Skill,适用于在华为 Ascend 910、910B、910C 上准备权重、适配 Lightning 和 CUDA only kernel、完成 Boltz2 端到端结构预测推理,并沉淀可复现的环境与验证命令。 | ascend-ai-coding/ | 174 | — | ~2.8k | Automated safety check: Pass | No licence | today |
| 267 | BoltzGen 昇腾 NPU 迁移与复现 Skill,适用于在华为 Ascend NPU 上部署 BoltzGen 生成式蛋白设计与逆折叠流程,覆盖环境准备、权重缓存、cuEquivariance 兼容、源码适配和端到端推理验证。 | ascend-ai-coding/ | 174 | — | ~3.8k | Automated safety check: Pass | No licence | today |
| 268 | DiffSBDD 昇腾 NPU 迁移 Skill,适用于将基于等变扩散模型的结构化药物设计项目从 CUDA 迁移到华为 Ascend NPU,覆盖环境搭建、依赖安装、torchscatter 源码编译、代码适配以及 de novo 推理验证。 | ascend-ai-coding/ | 174 | — | ~898 | Automated safety check: Pass | No licence | today |
| 269 | GENERator DNA 序列生成模型的昇腾 NPU 迁移 Skill,适用于将基于 HuggingFace Transformers 的 Causal LM 从 CUDA 迁移到华为 Ascend NPU,覆盖环境搭建、依赖安装、代码适配、多进程处理和 sequence recovery 验证。 | ascend-ai-coding/ | 174 | — | ~827 | Automated safety check: Pass | No licence | today |
| 270 | OligoFormer 昇腾 NPU 迁移 Skill,适用于将基于 PyTorch Transformer 的 siRNA 效能预测模型迁移到华为 Ascend NPU,覆盖环境搭建、RNA-FM 依赖安装、代码适配、推理验证与可选训练流程。 | ascend-ai-coding/ | 174 | — | ~1.3k | Automated safety check: Pass | No licence | today |
| 271 | 271.Comfy CLI Install, manage, and run ComfyUI instances. An agent skill from sundial-org/awesome-openclaw-skills. | sundial-org/ | 663 | — | ~1.5k | Automated safety check: Pass | No licence | 7 mo ago |
| 272 | 272.Deep Learning A skill your agent uses when training or debugging a neural net in PyTorch — the forward/loss/backward/step loop and its silent bugs, mixed precision (AMP), AdamW/LR schedules, DDP/FSDP/ZeRO… | ericrisco/ | 167 | — | ~3.4k | Automated safety check: Pass | MIT | today |
| 273 | 273.Cuda CUDA C/C++ skill for NVIDIA GPU kernel programming. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 274 | 274.Cuda Debugging CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 275 | 275.Cuda Profiling CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.6k | Automated safety check: Notes | MIT | 3 mo ago |
| 276 | 276.Hip Rocm HIP and ROCm skill for AMD GPU programming. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.6k | Automated safety check: Notes | MIT | 3 mo ago |
| 277 | 277.Triton Lang Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 278 | NVIDIA Collective Communications Library integration for multi-GPU operations. | majiayu000/ | 666 | 1 repo | ~1.9k | Automated safety check: Notes | MIT | today |
| 279 | 279.Unified Memory Expert skill for CUDA Unified Memory and memory prefetching optimization. | majiayu000/ | 666 | 1 repo | ~3.8k | Automated safety check: Notes | MIT | today |
| 280 | 280.Llama Cpp Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems. | magnus919/ | 113 | — | ~2.3k | Automated safety check: Pass | MIT | 2 days ago |
| 281 | 281.Nx Matting 使用本地 BiRefNet GGUF 模型完成图片或视频抠图、人物抠图、主体分割和背景移除,并输出透明 PNG、MOV 或 WebM。适用于用户提到图片抠图、照片去背景、人像透明图、视频抠图、透明视频、BiRefNet、JPG/PNG/BMP/WebP 图片,或 MP4/MOV/WebM 视频的场景;无需 Python、PyTorch 或 CUDA。 | aiskillstore/ | 430 | — | ~748 | Automated safety check: Pass | MIT | yesterday |
| 282 | 282.Uma Run structure relaxation and phonon calculations using Meta's UMA (Universal Materials Accelerator) via fairchem | lamm-mit/ | 244 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |