Search
CUDA · For data scientists
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | 1.Esmfold2 Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. | JimLiu/ | 228 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 2 | Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness. | stas00/ | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | 4 days ago |
| 3 | Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. | open-infra-skills/ | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 4 | A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing… | Mesh-LLM/ | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 5 | Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass. | KernelFlow-ops/ | 214 | — | ~4.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 6 | Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server. | huggingface/ | 11k | 3 repos | ~945 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 7 | Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator. | inclusionAI/ | 323 | — | ~498 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 8 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill. | mit-han-lab/ | 251 | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 10 | Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK). | intel/ | 1.6k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 12 | Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands. | brevdev/ | 146 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 13 | Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage | maxiaosong1124/ | 129 | — | ~1.6k | Automated safety check: Pass | MIT | 4 mo ago |
| 14 | Verify a CV-CUDA optimization campaign's deterministic definition-of-done and concise versioned MR summary per .agents/guidance/OPTIMIZATIONGUIDELINES.md. | CVCUDA/ | 2.7k | — | ~424 | Automated safety check: Pass | Unknown | 24 days ago |
| 15 | 15.Setup Guide A skill your agent uses when the user wants to set up LLM training for the first time, or when traininghub is not yet installed/configured in the current environment. | Red-Hat-AI-Innovation-Team/ | 100 | — | ~959 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 16 | 16.Model Deploy Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling). | sohu-mptc/ | 107 | — | ~974 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 17 | Add, fix, or validate Triton Runner support for an exact Triton version. | toyaix/ | 100 | — | ~1.1k | Automated safety check: Pass | MIT | 24 days ago |
| 18 | Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves… | Red-Hat-AI-Innovation-Team/ | 100 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 19 | Touch-lists for common OCANNL extension tasks: adding a primitive operation, adding or extending a backend, extending shape inference, and diagnosing output differences between backends. | ahrefs/ | 118 | — | ~728 | Automated safety check: Pass | BSD-2-Clause | 4 days ago |
| 20 | Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific… | alibaba/ | 168 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 21 | Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons. | Orchestra-Research/ | 13k | 2 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 22 | Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives. | Orchestra-Research/ | 13k | 4 repos | ~3k | Automated safety check: Warn | MIT | 3 mo ago |
| 23 | 23.Cuda Streams Manage tensor lifetimes across CUDA streams and decide whether recordstream is avoidable. | pytorch/ | 104k | — | ~796 | Automated safety check: Pass | Unknown | today |
| 24 | Add or debug an AReno model family, including config conversion, module construction, checkpoint load/save, text or multimodal inference, training backward, tensor parallelism, CUDA graph decode… | inclusionAI/ | 323 | — | ~722 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 25 | Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate). | CVCUDA/ | 2.7k | — | ~433 | Automated safety check: Pass | Unknown | 24 days ago |
| 26 | Review a CV-CUDA operator's BENCHMARK coverage — drivers, layout axis, baselines, the basic-tier floor, row counts, and coverage statistics. | CVCUDA/ | 2.7k | — | ~274 | Automated safety check: Pass | Unknown | 24 days ago |
| 27 | 27.GPU Backend pyqula's CPU/GPU switch (src/pyqula/gpu.py), how a routine is routed onto the device, per-call precision, and the tiered porting plan in documentation/gpuportingplan.md. | joselado/ | 145 | — | ~679 | Automated safety check: Pass | GPL-3.0 | 4 days ago |
| 28 | GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster. | K-Dense-AI/ | 48k | 1 repo | ~3.4k | Automated safety check: Pass | MIT | 6 days ago |
| 29 | Turns a parquet of image file paths into a parquet of embeddings with CLIP, SigLIP or a TAO checkpoint, using the TAO Data Services container, ahead of neighbor mining. | NVIDIA/ | 3.6k | — | ~2k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 30 | How to build and run GPU targets under Buck in fbcode. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~998 | Automated safety check: Pass | MIT | today |
| 31 | A skill your agent uses when the user needs to choose the tile, dtype, ctagroup::1/2, SMEM operand layout, or TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell, or to implement an… | mirage-project/ | 2.5k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 32 | Use FP16/BF16 mixed precision to accelerate training and reduce memory. | aiming-lab/ | 15k | — | ~275 | Automated safety check: Pass | MIT | 1 mo ago |
| 33 | 5-stage kernel correctness verification protocol for Triton and CUDA kernels. | ZJLi2013/ | 102 | — | ~702 | Automated safety check: Pass | No licence | 6 mo ago |
| 34 | A skill your agent uses when training, evaluating, or exporting Workflow policies with online RSL-RL or RLinf, including RL checkpoint and Workflow handoff. | NVIDIA/ | 3.6k | 1 repo | ~3.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 35 | Build deterministic forecast scripts with Earth2Studio (model, data source, IO, inference). | NVIDIA/ | 3.6k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 36 | NV-Tesseract AD Diffusion — diffusion-based anomaly detection and fine-tuning for multivariate time series. | NVIDIA/ | 3.6k | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 37 | NV-Tesseract Forecasting — transformer-based multivariate time series forecasting with DARR (context-enhanced kNN retrieval), interpretability, and fine-tuning. | NVIDIA/ | 3.6k | — | ~3.2k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 38 | Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent. | amd/ | 182 | — | ~4.8k | Automated safety check: Pass | MIT | 13 days ago |
| 39 | Guide installing Earth2Studio via uv or pip, selecting model extras, and configuring the environment. | NVIDIA/ | 3.6k | — | ~1.6k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 40 | Used for running NV-Segment-CTMR on CT or MRI NIfTI volumes and recording label-map evidence. | NVIDIA/ | 3.6k | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 41 | Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads. | NVIDIA/ | 3.6k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 42 | Generate inorganic material structures using MatterGen, a diffusion-based generative model. | learningmatter-mit/ | 176 | — | ~1.8k | Automated safety check: Pass | MIT | 3 days ago |
| 43 | GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. | Mathews-Tom/ | 329 | — | ~3.5k | Automated safety check: Notes | MIT | 5 days ago |
| 44 | 44.Make It 3D Use this repo skill for Make-It-3D single-image 3D creation, including CUDA asset setup, alpha-image validation, coarse NeRF optimization, refinement, rendering, export, and troubleshooting. | VectorSpaceLab/ | 331 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 45 | A skill your agent uses for Motional nuPlan autonomous-driving planning workflows: dataset and map access, scenario filtering, planner implementation, open- or closed-loop simulation, metrics… | VectorSpaceLab/ | 331 | — | ~1.2k | Automated safety check: Pass | Unknown | 1 mo ago |
| 46 | A skill your agent uses for Anomalib benchmark pipelines, tiled ensemble workflows, and advanced pipeline orchestration helpers while keeping experimental execution paths explicit. | VectorSpaceLab/ | 331 | — | ~952 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 47 | Use this repo skill for Microsoft Swin-Transformer image-classification model, config, data, checkpoint, SimMIM, Swin-MoE, and optional CUDA acceleration workflows. | VectorSpaceLab/ | 331 | — | ~1.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 48 | Generate novel crystal structures and molecules using ADiT (All-atom Diffusion Transformer), a unified latent diffusion model. | learningmatter-mit/ | 176 | — | ~1.4k | Automated safety check: Pass | MIT | 3 days ago |