AI model or service

vLLM agent skills, page 3

Skills #97–144 of 163, ranked by score.

vLLM skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

vLLM skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
97

Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…

ai-dynamo/dynamo8.2k—~4.9kAutomated safety check: PassApache-2.0today
98

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

vllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.06 mo ago
99

Given a .mlir file (or a directory of .mlir files) with TTIR ops, run the same TTIR normalization passes as D2MFrontendPipeline before D2M, then produce per-file outputs: preprocessed.mlir…

tenstorrent/tt-mlir311—~1.8kAutomated safety check: PassApache-2.0today
100

Select and verify the current region-specific serving container URI for a SageMaker model deployment.

waybarrios/opencode-power-pack533—~4.3kAutomated safety check: PassApache-2.0yesterday
101

Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL…

open-thoughts/OpenThoughts-Agent301—~1.2kAutomated safety check: PassApache-2.09 days ago
102

LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components.

ancoleman/ai-design-components5261 repo~3.4kAutomated safety check: PassMIT10 mo ago
103

Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.

sickn33/agentic-awesome-skills47k1 repo~1.9kAutomated safety check: PassApache-2.0yesterday
104

Curated upstream guidance for Huggingface Community Evals; use when the workflow matches the user goal.

sickn33/agentic-awesome-skills47k1 repo~1.7kAutomated safety check: PassMITyesterday
105

Build, test, and debug Hermes Agent RL environments for Atropos training.

Tommy-yw/RunbookHermes546—~3.3kAutomated safety check: PassMIT4 mo ago
106

Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio.

sickn33/agentic-awesome-skills47k2 repos~1.6kAutomated safety check: PassMITyesterday
107

Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

sickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMITyesterday
108

Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure.

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.5kAutomated safety check: PassApache-2.0today
109

Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation.

mirage-project/mirage2.5k—~4.6kAutomated safety check: PassApache-2.02 days ago
110

Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.

AMD-AGI/Hyperloom216—~7.2kAutomated safety check: NotesUnknowntoday
111
111.Vllm

Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.

Prism-Shadow/penguin-harness2.5k—~1kAutomated safety check: PassApache-2.0yesterday
112

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

wshobson/agents40k—~2kAutomated safety check: PassMIT3 days ago
113

Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard.

Xilinx/mlir-air150—~3.4kAutomated safety check: PassMITtoday
114

vLLM: high-throughput LLM serving, OpenAI API, quantization.

Luciole-Studio/Misaka-Agent1253 repos~2.3kAutomated safety check: PassMITtoday
115

Root-cause a vLLM torch-nightly CI regression report. An agent skill from pytorch/test-infra.

pytorch/test-infra113—~2.3kAutomated safety check: PassUnknowntoday
116

Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMITyesterday
117

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMITyesterday
118

Neutral, framework-agnostic decision tree for project kickoff: "which agent / RAG / LLM framework should I reach for?" Synthesizes the ecosystem sections of 7 landmark-project SOPs (LangGraph…

agentsope/SkillAlchemy457—~5.8kAutomated safety check: PassMIT1 mo ago
119

lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1252 repos~3.1kAutomated safety check: PassMITtoday
120

Review vLLM-routed torch-nightly root causes against existing upstream vLLM issues using read-only search results, and emit a validated upstream-checks artifact for the filer.

pytorch/test-infra113—~1.3kAutomated safety check: PassUnknowntoday
121

Launch a datagen (trace-generation) job on an HPC cluster (Jupiter/Leonardo/Perlmutter) via the OpenThoughts-Agent hpc.launch --jobtype datagen entrypoint — the cluster-AGNOSTIC general flow…

open-thoughts/OpenThoughts-Agent301—~2.5kAutomated safety check: PassApache-2.09 days ago
122

Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a…

ruvnet/ruflo74k—~336Automated safety check: NotesMITtoday
123

[omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the…

rlaope/oh-my-hermes3.2k—~2.2kAutomated safety check: PassMITtoday
124

Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.

open-thoughts/OpenThoughts-Agent301—~947Automated safety check: PassApache-2.09 days ago
125

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…

open-thoughts/OpenThoughts-Agent301—~3.7kAutomated safety check: PassApache-2.09 days ago
126

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

agentsope/SkillAlchemy457—~6.1kAutomated safety check: PassMIT1 mo ago
127

Project-kickoff rubric for the self-host vs managed-cloud decision — when is running your own inference engine / LLM platform worth the ops cost vs paying per-token for a managed API?

agentsope/SkillAlchemy457—~6.3kAutomated safety check: PassMIT1 mo ago
128

Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy.

agentsope/SkillAlchemy457—~6.1kAutomated safety check: PassMIT1 mo ago
129

Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…

autonomous-ai/openharness1.1k—~1.9kAutomated safety check: PassMITtoday
130

Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent.

open-thoughts/OpenThoughts-Agent301—~5.8kAutomated safety check: PassApache-2.09 days ago
131

EXECUTE a staged codebase plan (from code-create-staged-plan or an existing notes/<codebase/ plan) one stage at a time, gate-by-gate, while keeping the local clone ground truth and a dated…

open-thoughts/OpenThoughts-Agent301—~977Automated safety check: PassApache-2.09 days ago
132

Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics.

benchflow-ai/skillsbench1.8k—~2.3kAutomated safety check: PassApache-2.02 mo ago
133

Add and manage evaluation results in Hugging Face model cards.

majiayu000/claude-skill-registry6663 repos~5.6kAutomated safety check: NotesMITtoday
134

End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks.

amd/Quark181—~6.2kAutomated safety check: PassMIT10 days ago
135

This skill should be used when the user asks about "local models", "custom models", "fine-tuning", "self-hosting models", "model selection", "which model should I use", "data privacy and models"…

Habitat-Thinking/ai-literacy-superpowers114—~1kAutomated safety check: PassUnknown17 days ago
136

vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licencetoday
137

Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode…

ascend-ai-coding/awesome-ascend-skills174—~731Automated safety check: PassNo licencetoday
138
138.Vllm

A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…

majiayu000/claude-skill-registry6661 repo~3.6kAutomated safety check: PassMITtoday
139
139.Vllm

vLLM is a high-throughput inference and serving engine for large language models that exposes an OpenAI-compatible HTTP API and a Python batch API.

majiayu000/claude-skill-registry6661 repo~1.9kAutomated safety check: PassApache-2.0today
140

Operate, configure, secure, and troubleshoot the LiteLLM AI gateway (proxy) and Python SDK: run the proxy (litellm --config), route to 100+ providers through one OpenAI-compatible API, configure…

magnus919/agent-skills111—~4.2kAutomated safety check: NotesMITyesterday
141
141.Vllm

Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

magnus919/agent-skills111—~4.1kAutomated safety check: NotesMITyesterday
142

Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve.

ascend-ai-coding/awesome-ascend-skills174—~5.6kAutomated safety check: PassNo licencetoday
143

Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

AnastasiyaW/codex-claude-code-config154—~794Automated safety check: PassMIT5 days ago
144

Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

BagelHole/DevOps-Security-Agent-Skills1.1k—~2kAutomated safety check: PassMIT4 mo ago