Search

AI & LLM Engineering · CUDA

209 skills found, page 2.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

Register, extend, or audit instructions in the table-driven T.ptx dialect (python/tvm/backend/cuda/ptx/table.py), and move the table to a newer PTX ISA version.

mlc-ai/relax175—~5.1kAutomated safety check: PassApache-2.07 days ago
50

Scaffold new Celeritas source and test files with the required copyright header and register them in CMake.

celeritas-project/celeritas105—~496Automated safety check: PassUnknowntoday
51

Verify a CV-CUDA optimization campaign's deterministic definition-of-done and concise versioned MR summary per .agents/guidance/OPTIMIZATIONGUIDELINES.md.

CVCUDA/CV-CUDA2.7k—~424Automated safety check: PassUnknown24 days ago
52

A skill your agent uses when the user wants to set up LLM training for the first time, or when traininghub is not yet installed/configured in the current environment.

Red-Hat-AI-Innovation-Team/training_hub100—~959Automated safety check: PassApache-2.04 days ago
53

Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration.

vllm-project/vllm-omni7.1k—~8.7kAutomated safety check: PassApache-2.0today
54

Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling).

sohu-mptc/FlashRec107—~974Automated safety check: PassApache-2.0yesterday
55

Guided workflow for adding a new model architecture to llama.cpp.

JakeATX/llamAmpere166—~4.1kAutomated safety check: PassMITyesterday
56

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…

vipshop/cache-dit1.3k—~2.2kAutomated safety check: PassApache-2.0yesterday
57

Add, fix, or validate Triton Runner support for an exact Triton version.

toyaix/triton-runner100—~1.1kAutomated safety check: PassMIT24 days ago
58

A skill your agent uses for performance profiling and optimization.

ByteDance-Seed/VeOmni2.2k—~1.7kAutomated safety check: PassApache-2.0yesterday
59

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware.

Orchestra-Research/AI-Research-SKILLs13k3 repos~1.5kAutomated safety check: PassMIT3 mo ago
60

Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed…

pytorch/test-infra113—~2kAutomated safety check: PassUnknowntoday
61
61.Compileiq DebugOfficial

A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is…

NVIDIA/CompileIQ138—~2.9kAutomated safety check: NotesApache-2.018 days ago
62

Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits.

huggingface/skills11k2 repos~4.6kAutomated safety check: PassApache-2.03 days ago
63

Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2kAutomated safety check: PassNo licence6 days ago
64

Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)

sgl-project/sglang37k2 repos~3.4kAutomated safety check: PassApache-2.0today
65

Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging

sgl-project/sglang37k2 repos~4.9kAutomated safety check: PassApache-2.0today
66

Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves…

Red-Hat-AI-Innovation-Team/training_hub100—~2.8kAutomated safety check: PassApache-2.04 days ago
67

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

slowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT2 mo ago
68

Add a new CV-CUDA operator end-to-end per .agents/guidance/MAKEOPGUIDELINES.md, with a deterministically-enforced definition-of-done.

CVCUDA/CV-CUDA2.7k—~831Automated safety check: PassUnknown24 days ago
69

A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing…

vipshop/cache-dit1.3k—~3.8kAutomated safety check: PassApache-2.0yesterday
70

Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI.

lucifer1004/VeloQ128—~1.3kAutomated safety check: PassMIT6 days ago
71
71.Cuda

Draft, debug, and measure CUDA kernels and host launch workflows.

sablin39/tilelang-cuda-skills145—~990Automated safety check: PassNo licence26 days ago
72

Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired…

NVIDIA/cosmos-framework560—~2.7kAutomated safety check: PassUnknownyesterday
73

Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing.

microsoft/onnxruntime22k—~6.5kAutomated safety check: PassMITtoday
74

Touch-lists for common OCANNL extension tasks: adding a primitive operation, adding or extending a backend, extending shape inference, and diagnosing output differences between backends.

ahrefs/ocannl118—~728Automated safety check: PassBSD-2-Clause4 days ago
75
75.At Dispatch V2Official

Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

intel/torch-xpu-ops1153 repos~2.2kAutomated safety check: PassApache-2.0yesterday
76

A skill your agent uses when composing the Search(...) call and calling .start().

NVIDIA/CompileIQ138—~2.2kAutomated safety check: NotesApache-2.018 days ago
77

Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder).

matlab/agent-skills-playground184—~3.4kAutomated safety check: PassUnknownyesterday
78

Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.

JakeATX/llamAmpere166—~5.6kAutomated safety check: PassMITyesterday
79

Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific…

alibaba/atrex-kernel-agent168—~1.3kAutomated safety check: PassApache-2.0yesterday
80

Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works.

vllm-project/vllm-omni7.1k—~5.5kAutomated safety check: PassApache-2.0today
81

Review 3DGS implementation code for correctness, performance bugs, and best practices.

jaccen/Awesome-Gaussian-Skills161—~2.9kAutomated safety check: PassApache-2.0today
82

使用 PaddlePaddle 仓库规则评审 Pull Request 和全仓库代码变更,覆盖正确性、兼容性、算子、分布式、数值、性能、安全、测试、构建和 PR 信息。当需要审查 Paddle 的代码、测试、算子 YAML、C++/CUDA/XPU kernel、Python API、分布式逻辑或 CI 配置时使用。

PaddlePaddle/Paddle24k—~303Automated safety check: PassApache-2.0yesterday
83

Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

Orchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT3 mo ago
84

Scaffold a new CV-CUDA operator — a complete, wired, building skeleton — and delegate the implementation to a human or another AI.

CVCUDA/CV-CUDA2.7k—~306Automated safety check: PassUnknown24 days ago
85

Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

facebookexperimental/triton201—~1.6kAutomated safety check: PassMITtoday
86

Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.

Orchestra-Research/AI-Research-SKILLs13k4 repos~3kAutomated safety check: WarnMIT3 mo ago
87

Manage tensor lifetimes across CUDA streams and decide whether recordstream is avoidable.

pytorch/pytorch104k—~796Automated safety check: PassUnknowntoday
88

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

vllm-project/vllm-skills102—~1.6kAutomated safety check: PassApache-2.06 mo ago
89

Workflow for upgrading/integrating cache-dit in SGLang diffusion (multimodalgen): DBCache, DMD calibrator, TaylorSeer, SVDQuant DQ; porting upstream PRs and resolving conflicts against the…

sgl-project/sglang37k—~7.1kAutomated safety check: PassApache-2.0today
90

Rigor Debug / Rigor Audit skill for deep learning research work.

lllllllama/RigorPilot-Skills4971 repo~522Automated safety check: PassMIT18 days ago
91

Add or debug an AReno model family, including config conversion, module construction, checkpoint load/save, text or multimodal inference, training backward, tensor parallelism, CUDA graph decode…

inclusionAI/AReno323—~722Automated safety check: PassApache-2.0yesterday
92

Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate).

CVCUDA/CV-CUDA2.7k—~433Automated safety check: PassUnknown24 days ago
93

Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

facebookexperimental/triton201—~709Automated safety check: PassMITtoday
94

Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

amd/Quark182—~1.4kAutomated safety check: PassMIT13 days ago
95

Update .vscode/launch.json to debug a specific CTest test by name.

celeritas-project/celeritas105—~199Automated safety check: PassUnknowntoday
96

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills408—~2.3kAutomated safety check: PassMIT2 days ago