Search

NVIDIA AI Platform · LLM inference and serving

55 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

sgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0today
2

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

dstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.02 days ago
3

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.02 days ago
4

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNo licence6 days ago
5

Integrates a hosted API or local/open-weight model end-to-end across OpenTryOn (adapter, CLI registry, MCP, docs) and TryOn Studio (catalog, Connect keys, planner).

tryonlabs/opentryon551—~1.1kAutomated safety check: PassUnknown2 days ago
6

Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable.

EGalahad/sim2real146—~1.1kAutomated safety check: PassNo licence13 days ago
7

Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.

dstackai/dstack2.3k—~403Automated safety check: PassMPL-2.02 days ago
8

Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data, recipes, evaluation, quantization, deployment.

NVIDIA-NeMo/Nemotron2.1k—~1.9kAutomated safety check: PassApache-2.05 days ago
9

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

Orchestra-Research/AI-Research-SKILLs13k4 repos~1.3kAutomated safety check: PassMIT3 mo ago
10

Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

amd/skills408—~4kAutomated safety check: NotesMIT2 days ago
11

Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes.

NVIDIA-NeMo/Nemotron2.1k—~2.4kAutomated safety check: PassApache-2.05 days ago
12

GGUF format and llama.cpp quantization for efficient CPU/GPU inference.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.6kAutomated safety check: PassMIT3 mo ago
13

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware.

Orchestra-Research/AI-Research-SKILLs13k3 repos~1.5kAutomated safety check: PassMIT3 mo ago
14

Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.

scragnog/HOT-Step-CPP174—~4.9kAutomated safety check: PassMIT2 days ago
15

Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
16

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

vllm-project/vllm-skills102—~1.6kAutomated safety check: PassApache-2.06 mo ago
17

Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.3kAutomated safety check: PassMIT3 mo ago
18

Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

amd/Quark182—~1.4kAutomated safety check: PassMIT13 days ago
19

Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference.

NVIDIA-NeMo/Nemotron2.1k—~1.8kAutomated safety check: PassApache-2.05 days ago
20

A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~1.5kAutomated safety check: PassNo licence6 days ago
21

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

vllm-project/vllm-skills102—~2.5kAutomated safety check: NotesApache-2.06 mo ago
22

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
23

LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components.

ancoleman/ai-design-components525—~3.4kAutomated safety check: PassMIT10 mo ago
24
24.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.6k1 repo~1.8kAutomated safety check: PassApache-2.02 days ago
25

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

sickn33/agentic-awesome-skills47k1 repo~2.3kAutomated safety check: PassMIT2 days ago
26

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

NVIDIA/skills3.6k1 repo~2.9kAutomated safety check: PassApache-2.02 days ago
27

Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

NVIDIA/skills3.6k1 repo~3.1kAutomated safety check: PassApache-2.02 days ago
28
28.Gke InferenceOfficial

Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.

google/skills21k—~2kAutomated safety check: PassApache-2.0yesterday
29

High-throughput LLM inference on NVIDIA GPUs. An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1711 repo~1.3kAutomated safety check: PassMIT3 days ago
30

Automatically fetch InferenceX benchmark data and generate daily performance reports for LLM inference on various hardware (NVIDIA, AMD, etc.).

ascend-ai-coding/awesome-ascend-skills174—~1.1kAutomated safety check: PassNo licenceyesterday
31
31.Jetson LLM ServeOfficial

Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.

NVIDIA/skills3.6k1 repo~3kAutomated safety check: NotesApache-2.02 days ago
32

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

NVIDIA/skills3.6k1 repo~1.2kAutomated safety check: PassApache-2.02 days ago
33
33.Deepstream DevOfficial

NVIDIA DeepStream SDK development with Python pyservicemaker API.

NVIDIA/skills3.6k—~3.3kAutomated safety check: PassApache-2.02 days ago
34
34.Deepstream SopOfficial

A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether…

NVIDIA/skills3.6k—~4.7kAutomated safety check: NotesApache-2.02 days ago
35

Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines.

NVIDIA/skills3.6k—~1.5kAutomated safety check: PassApache-2.02 days ago
36

A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill.

NVIDIA/skills3.6k—~2.1kAutomated safety check: PassApache-2.02 days ago
37

CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

NVIDIA/skills3.6k—~4kAutomated safety check: NotesApache-2.02 days ago
38

InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment.

NVIDIA/skills3.6k—~3.5kAutomated safety check: NotesApache-2.02 days ago
39

The mandatory pre-launch gate and four-verb execution contract for every TAO workflow or action.

NVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.02 days ago
40

PyTorch-based TAO image classification. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k—~3.6kAutomated safety check: NotesApache-2.02 days ago
41
41.Tao Train RtdetrOfficial

RT-DETR (Real-Time DEtection TRansformer) for 2D object detection.

NVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.02 days ago
42

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

agentsope/SkillAlchemy436—~6.1kAutomated safety check: PassMIT2 days ago
43

Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy.

agentsope/SkillAlchemy436—~6.1kAutomated safety check: PassMIT2 days ago
44

Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…

autonomous-ai/openharness1.2k—~1.9kAutomated safety check: PassMITyesterday
45

Diagnose Day-2 AKS GPU and KAITO incidents using profile-aware, read-only evidence.

microsoft/GitHub-Copilot-for-Azure255—~764Automated safety check: PassMITyesterday
46

How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.

NVIDIA/skills3.6k—~5kAutomated safety check: NotesApache-2.02 days ago
47

Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).

NVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.02 days ago
48

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: NotesMITyesterday