Search

vLLM · LLM inference and serving

145 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

sgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0today
2

Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

huggingface/skills11k1 repo~4.6kAutomated safety check: PassApache-2.02 days ago
3

Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

amElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT3 days ago
4

Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

huggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.02 days ago
5

Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

vllm-project/vllm-omni7.1k—~7.5kAutomated safety check: PassApache-2.0today
6

Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

vllm-project/vllm-ascend2.9k—~7.2kAutomated safety check: PassApache-2.0today
7

Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

guqiong96/Lvllm4652 repos~349Automated safety check: PassApache-2.019 days ago
8

Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

huggingface/skills11k1 repo~6.9kAutomated safety check: PassApache-2.02 days ago
9

Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

vllm-project/vllm-omni7.1k—~1.4kAutomated safety check: PassApache-2.0today
10

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

dstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.02 days ago
11

Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

guqiong96/Lvllm4652 repos~1.5kAutomated safety check: PassApache-2.019 days ago
12
12.Debug InferenceOfficial

Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.

NVIDIA/OpenShell16k—~1.9kAutomated safety check: PassApache-2.0yesterday
13

Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2.5kAutomated safety check: PassNo licence6 days ago
14

Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

vllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0today
15

Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings.

vllm-project/vllm-omni7.1k—~3.8kAutomated safety check: PassApache-2.0today
16

Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

Blackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT3 days ago
17

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0yesterday
18

驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

OpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.01 mo ago
19

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNo licence6 days ago
20

Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

ModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassUnknowntoday
21

Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).

intel/auto-round1.6k—~1.9kAutomated safety check: PassApache-2.0yesterday
22

Adapt and port new LLM model architectures to this xinfer project.

guoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT1 mo ago
23

Creates and validates dubbo-go-pixiu LLM gateway conf.yaml. An agent skill from apache/dubbo-go-pixiu.

apache/dubbo-go-pixiu567—~2.5kAutomated safety check: PassApache-2.08 days ago
24

Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

vllm-project/vllm-omni7.1k—~7kAutomated safety check: PassApache-2.0today
25

Always resolve Hugging Face models via model-shelf before any download.

alexziskind1/model-shelf130—~792Automated safety check: PassMIT1 mo ago
26

Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

MetaX-MACA/vLLM-metax180—~3.2kAutomated safety check: PassApache-2.0yesterday
27

Build the standalone ZenDNN native library (zendnnl) from source with the alternate compute backends OFF (no oneDNN, libxsmm, parlooper, fbgemm), keeping AOCL DLP, which is the GEMM backend zendnnl…

amd/ZenDNN158—~2kAutomated safety check: PassUnknown4 days ago
28

Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

guoqingbao/xinfer334—~3.8kAutomated safety check: PassMIT1 mo ago
29

Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models.

Orchestra-Research/AI-Research-SKILLs13k9 repos~4kAutomated safety check: PassMIT3 mo ago
30

Add or update an in-repository vLLM-Omni model recipe with verified task, input, output, hardware, command, feature, and validation contracts.

vllm-project/vllm-omni7.1k—~1.4kAutomated safety check: PassApache-2.0today
31

Review and refactor inference-runtime code using concrete rules for responsibility boundaries, state ownership, interfaces, asynchronous lifetimes, KV management, and maintainability.

ThinkFlowLab/vllm-rlt149—~4.2kAutomated safety check: PassApache-2.0today
32

Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

Orchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT3 mo ago
33

Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.

guqiong96/Lvllm4651 repo~831Automated safety check: PassApache-2.019 days ago
34

Review and adapt vllmmetax/registry registrations, quantization configurations, CustomOps and kernel dispatch against a target vLLM revision and installed MetaX APIs.

MetaX-MACA/vLLM-metax180—~1.8kAutomated safety check: PassApache-2.0yesterday
35

Advisory guidance for Amazon EKS architecture and configuration decisions — compute strategy, networking, security, reliability, cost, autoscaling, observability, multi-tenancy, and upgrade planning.

aws-samples/appmod-blueprints115—~5kAutomated safety check: PassMIT-02 days ago
36

Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding.

marin-community/marin3.9k—~1.1kAutomated safety check: PassApache-2.0today
37

Test LLM models served by xinfer for correctness, output quality, and performance.

guoqingbao/xinfer334—~2.6kAutomated safety check: PassMIT1 mo ago
38

Validate changed quantization math, packed formats, and inference kernels against a Torch oracle; adjudicate near-tie mismatches before rejecting optimized arithmetic.

ModelCloud/GPTQModel1.3k—~1.4kAutomated safety check: PassUnknowntoday
39

Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

vllm-project/vllm-skills102—~1.5kAutomated safety check: PassApache-2.06 mo ago
40

Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

amd/skills408—~4kAutomated safety check: NotesMITyesterday
41
41.Aqua DeploymentOfficial

Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.

oracle/accelerated-data-science125—~2.4kAutomated safety check: PassUPL-1.01 mo ago
42

Build vLLM (CPU) end-to-end with the oneDNN + ZenDNN (zen64) backend via direct integration (NO zentorch).

amd/ZenDNN158—~5.2kAutomated safety check: PassUnknown4 days ago
43

Review PRs and local changes for hsliuustc0106/vllm-rlt: Ouro engine and KV correctness, serving behavior, and BF16 accuracy/speed evidence.

ThinkFlowLab/vllm-rlt149—~1.1kAutomated safety check: PassApache-2.0today
44

Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

agentsope/SkillAlchemy436—~3kAutomated safety check: PassMIT2 days ago
45

Find evidence-backed simplification candidates in vLLM-Omni and, when requested, turn them into focused proposals or code changes.

vllm-project/vllm-omni7.1k—~2.7kAutomated safety check: PassApache-2.0today
46

Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.

sohu-mptc/FlashRec107—~915Automated safety check: PassApache-2.0yesterday
47

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

vllm-project/vllm-skills102—~1.6kAutomated safety check: PassApache-2.06 mo ago
48

Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

amd/Quark182—~3kAutomated safety check: PassMIT13 days ago