Search

vLLM

173 skills found, page 3.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
97

Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags.

waybarrios/opencode-power-pack534—~4.6kAutomated safety check: PassApache-2.05 days ago
98

Diagnose and fix OCI AI Quick Actions (AQUA) issues including deployment failures, OOM errors, authorization problems, capacity issues, container errors, and policy misconfigurations.

oracle/accelerated-data-science125—~1.8kAutomated safety check: PassUPL-1.01 mo ago
99

Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.

vllm-project/vllm-omni7.1k—~3kAutomated safety check: PassApache-2.0today
100

Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…

ai-dynamo/dynamo8.3k—~4.9kAutomated safety check: PassApache-2.0today
101

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

vllm-project/vllm-skills102—~2.5kAutomated safety check: NotesApache-2.06 mo ago
102

Given a .mlir file (or a directory of .mlir files) with TTIR ops, run the same TTIR normalization passes as D2MFrontendPipeline before D2M, then produce per-file outputs: preprocessed.mlir…

tenstorrent/tt-mlir314—~1.8kAutomated safety check: PassApache-2.0yesterday
103

Select and verify the current region-specific serving container URI for a SageMaker model deployment.

waybarrios/opencode-power-pack534—~4.3kAutomated safety check: PassApache-2.05 days ago
104

Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL…

open-thoughts/OpenThoughts-Agent301—~1.2kAutomated safety check: PassApache-2.012 days ago
105

Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.

sickn33/agentic-awesome-skills47k1 repo~1.9kAutomated safety check: PassApache-2.0yesterday
106

Curated upstream guidance for Huggingface Community Evals; use when the workflow matches the user goal.

sickn33/agentic-awesome-skills47k1 repo~1.7kAutomated safety check: PassMITyesterday
107

Build, test, and debug Hermes Agent RL environments for Atropos training.

Tommy-yw/RunbookHermes546—~3.3kAutomated safety check: PassMIT4 mo ago
108

Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio.

sickn33/agentic-awesome-skills47k2 repos~1.6kAutomated safety check: PassMIT2 days ago
109

Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

sickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT2 days ago
110

Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure.

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.5kAutomated safety check: PassApache-2.0yesterday
111

Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation.

mirage-project/mirage2.5k—~4.6kAutomated safety check: PassApache-2.03 days ago
112

Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.

AMD-AGI/Hyperloom219—~7.1kAutomated safety check: NotesUnknownyesterday
113
113.Vllm

Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.

Prism-Shadow/penguin-harness2.5k—~1kAutomated safety check: PassApache-2.0yesterday
114

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
115

Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard.

Xilinx/mlir-air150—~3.4kAutomated safety check: PassMITyesterday
116

LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components.

ancoleman/ai-design-components525—~3.4kAutomated safety check: PassMIT10 mo ago
117

Root-cause a vLLM torch-nightly CI regression report. An agent skill from pytorch/test-infra.

pytorch/test-infra113—~2.4kAutomated safety check: PassUnknownyesterday
118

Add and manage evaluation results in Hugging Face model cards.

sickn33/agentic-awesome-skills47k2 repos~418Automated safety check: PassMIT2 days ago
119

Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data.

NVIDIA/skills3.6k1 repo~2.3kAutomated safety check: NotesApache-2.0yesterday
120
120.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.6k1 repo~1.8kAutomated safety check: PassApache-2.0yesterday
121

Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMIT2 days ago
122

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMIT2 days ago
123

vLLM: high-throughput LLM serving, OpenAI API, quantization.

Luciole-Studio/Misaka-Agent1712 repos~2.3kAutomated safety check: PassMIT3 days ago
124

lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1712 repos~3.1kAutomated safety check: PassMIT3 days ago
125

Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters.

google/skills21k—~3.1kAutomated safety check: PassApache-2.0yesterday
126

Neutral, framework-agnostic decision tree for project kickoff: "which agent / RAG / LLM framework should I reach for?" Synthesizes the ecosystem sections of 7 landmark-project SOPs (LangGraph…

agentsope/SkillAlchemy436—~5.8kAutomated safety check: PassMIT2 days ago
127

Review vLLM-routed torch-nightly root causes against existing upstream vLLM issues using read-only search results, and emit a validated upstream-checks artifact for the filer.

pytorch/test-infra113—~1.3kAutomated safety check: PassUnknownyesterday
128

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

NVIDIA/skills3.6k1 repo~2.9kAutomated safety check: PassApache-2.0yesterday
129

Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

NVIDIA/skills3.6k1 repo~3.1kAutomated safety check: PassApache-2.0yesterday
130

Launch a datagen (trace-generation) job on an HPC cluster (Jupiter/Leonardo/Perlmutter) via the OpenThoughts-Agent hpc.launch --jobtype datagen entrypoint — the cluster-AGNOSTIC general flow…

open-thoughts/OpenThoughts-Agent301—~2.5kAutomated safety check: PassApache-2.012 days ago
131
131.ML AIOfficial

Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN…

grafana/skills282—~1.3kAutomated safety check: PassApache-2.02 days ago
132

Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a…

ruvnet/ruflo74k—~336Automated safety check: NotesMITyesterday
133
133.Airunway Aks SetupOfficial

Set up AI Runway on AKS — from bare cluster to running model.

microsoft/GitHub-Copilot-for-Azure2551 repo~1.1kAutomated safety check: PassMITyesterday
134

[omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the…

rlaope/oh-my-hermes3.2k—~2.2kAutomated safety check: PassMITyesterday
135
135.Jetson LLM ServeOfficial

Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.

NVIDIA/skills3.6k1 repo~3kAutomated safety check: NotesApache-2.0yesterday
136

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

NVIDIA/skills3.6k1 repo~1.2kAutomated safety check: PassApache-2.0yesterday
137

Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.

open-thoughts/OpenThoughts-Agent301—~947Automated safety check: PassApache-2.012 days ago
138

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…

open-thoughts/OpenThoughts-Agent301—~3.7kAutomated safety check: PassApache-2.012 days ago
139
139.Deepstream SopOfficial

A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether…

NVIDIA/skills3.6k—~4.7kAutomated safety check: NotesApache-2.0yesterday
140

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

agentsope/SkillAlchemy436—~6.1kAutomated safety check: PassMIT2 days ago
141

Project-kickoff rubric for the self-host vs managed-cloud decision — when is running your own inference engine / LLM platform worth the ops cost vs paying per-token for a managed API?

agentsope/SkillAlchemy436—~6.3kAutomated safety check: PassMIT2 days ago
142

Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy.

agentsope/SkillAlchemy436—~6.1kAutomated safety check: PassMIT2 days ago
143

Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…

autonomous-ai/openharness1.2k—~1.9kAutomated safety check: PassMITyesterday
144

Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent.

open-thoughts/OpenThoughts-Agent301—~5.8kAutomated safety check: PassApache-2.012 days ago