Search

CUDA · LLM inference and serving

31 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

guqiong96/Lvllm4652 repos~1.5kAutomated safety check: PassApache-2.019 days ago
2

Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

Blackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT3 days ago
3

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0yesterday
4

A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

Mesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0today
5

Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

huggingface/skills11k3 repos~945Automated safety check: PassApache-2.02 days ago
6

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNo licence6 days ago
7

Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).

intel/auto-round1.6k—~2.3kAutomated safety check: PassApache-2.0today
8

Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).

zhongkaifu/TensorSharp568—~1kAutomated safety check: PassBSD-3-Clausetoday
9

Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better.

Aseiel/VideoHighlighter166—~839Automated safety check: PassAGPL-3.0today
10

A skill your agent uses for web searches and current information lookups, finding sources, fact-checking, researching questions, comparing sources, or summarising web pages.

zhongkaifu/TensorSharp568—~2.3kAutomated safety check: WarnBSD-3-Clausetoday
11

Guided workflow for adding a new model architecture to llama.cpp.

JakeATX/llamAmpere166—~4.1kAutomated safety check: PassMITyesterday
12

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware.

Orchestra-Research/AI-Research-SKILLs13k3 repos~1.5kAutomated safety check: PassMIT3 mo ago
13

Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.

JakeATX/llamAmpere166—~5.6kAutomated safety check: PassMITyesterday
14

Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works.

vllm-project/vllm-omni7.1k—~5.5kAutomated safety check: PassApache-2.0today
15

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

vllm-project/vllm-skills102—~1.6kAutomated safety check: PassApache-2.06 mo ago
16

Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

amd/Quark182—~1.4kAutomated safety check: PassMIT13 days ago
17

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills408—~2.3kAutomated safety check: PassMITyesterday
18
18.App

Opinionated app components building on top of ./ui primitives

JakeATX/llamAmpere166—~146Automated safety check: PassMITyesterday
19

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

vllm-project/vllm-skills102—~2.5kAutomated safety check: NotesApache-2.06 mo ago
20

Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

sickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT2 days ago
21

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
22
22.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.6k1 repo~1.8kAutomated safety check: PassApache-2.0yesterday
23

Inspect a target ONNX model and prepare metadata for Quark ONNX PTQ planning.

amd/Quark182—~4.3kAutomated safety check: PassMIT13 days ago
24
24.Deepstream DevOfficial

NVIDIA DeepStream SDK development with Python pyservicemaker API.

NVIDIA/skills3.6k—~3.3kAutomated safety check: PassApache-2.0yesterday
25

Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines.

NVIDIA/skills3.6k—~1.5kAutomated safety check: PassApache-2.0yesterday
26

A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill.

NVIDIA/skills3.6k—~2.1kAutomated safety check: PassApache-2.0yesterday
27

Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent.

amd/Quark182—~4.8kAutomated safety check: PassMIT13 days ago
28

Diagnose failed Quark installation, PTQ execution, script generation, or export attempts.

amd/Quark182—~1.9kAutomated safety check: NotesMIT13 days ago
29

Use this sub-skill for Torch-TensorRT runtime performance controls, CUDA Graphs, output allocation, caches, TensorRT-RTX runtime settings, mutable modules, refit, weight streaming, and benchmark…

VectorSpaceLab/AREX-Skill331—~1kAutomated safety check: PassBSD-3-Clause1 mo ago
30

A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging…

VectorSpaceLab/AREX-Skill331—~1.5kAutomated safety check: PassBSD-3-Clause1 mo ago
31

Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems.

magnus919/agent-skills115—~2.3kAutomated safety check: PassMITyesterday