Search
vLLM
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints. | Orchestra-Research/ | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | 3 mo ago |
| 2 | Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed. | sgl-project/ | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones. | huggingface/ | 11k | 1 repo | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 4 | Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself. | amElnagdy/ | 2.3k | 2 repos | ~3k | Automated safety check: Pass | MIT | 3 days ago |
| 5 | Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends. | huggingface/ | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 6 | Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation. | vllm-project/ | 7.1k | — | ~7.5k | Automated safety check: Pass | Apache-2.0 | today |
| 7 | Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements. | vllm-project/ | 2.9k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | today |
| 8 | Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness. | vllm-project/ | 7.1k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm. | guqiong96/ | 465 | 2 repos | ~349 | Automated safety check: Pass | Apache-2.0 | 18 days ago |
| 10 | Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups. | huggingface/ | 11k | 1 repo | ~6.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 11 | 11.Quantization Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models. | vllm-project/ | 7.1k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 12 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | yesterday |
| 13 | Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL… | guqiong96/ | 465 | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | 18 days ago |
| 14 | Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM. | NVIDIA/ | 16k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 15 | Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths. | BBuf/ | 938 | — | ~2.5k | Automated safety check: Pass | No licence | 5 days ago |
| 16 | Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit. | vllm-project/ | 2.9k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 17 | 17.Review PR Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings. | vllm-project/ | 7.1k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 18 | Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks. | Blackwellboy/ | 135 | — | ~2.1k | Automated safety check: Pass | MIT | 2 days ago |
| 19 | 19.Graphsignal Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | today |
| 20 | 20.One Eval 驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。 | OpenDCAI/ | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 21 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 938 | — | ~2.8k | Automated safety check: Pass | No licence | 5 days ago |
| 22 | Low-token Codex session/thread title organizer. An agent skill from David-Lzy/codex_session_renamer. | David-Lzy/ | 108 | — | ~1.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 23 | Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems. | ModelCloud/ | 1.3k | — | ~1.1k | Automated safety check: Pass | Unknown | today |
| 24 | Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor). | intel/ | 1.6k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 25 | 25.Add Model Adapt and port new LLM model architectures to this xinfer project. | guoqingbao/ | 334 | — | ~4.2k | Automated safety check: Notes | MIT | 1 mo ago |
| 26 | Creates and validates dubbo-go-pixiu LLM gateway conf.yaml. An agent skill from apache/dubbo-go-pixiu. | apache/ | 567 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 27 | Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT… | vllm-project/ | 7.1k | — | ~7k | Automated safety check: Pass | Apache-2.0 | today |
| 28 | Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 5 days ago |
| 29 | 29.Resolve Always resolve Hugging Face models via model-shelf before any download. | alexziskind1/ | 130 | — | ~792 | Automated safety check: Pass | MIT | 1 mo ago |
| 30 | Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels. | MetaX-MACA/ | 180 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | today |
| 31 | 31.Build Zendnn Build the standalone ZenDNN native library (zendnnl) from source with the alternate compute backends OFF (no oneDNN, libxsmm, parlooper, fbgemm), keeping AOCL DLP, which is the GEMM backend zendnnl… | amd/ | 158 | — | ~2k | Automated safety check: Pass | Unknown | 3 days ago |
| 32 | 32.Check Model Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer. | guoqingbao/ | 334 | — | ~3.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 33 | Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models. | Orchestra-Research/ | 13k | 9 repos | ~4k | Automated safety check: Pass | MIT | 3 mo ago |
| 34 | 34.Add Recipe Add or update an in-repository vLLM-Omni model recipe with verified task, input, output, hardware, command, feature, and validation contracts. | vllm-project/ | 7.1k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 35 | Quick-reference help for fine-tuning language models with Axolotl, covering YAML configs, FSDP, context parallelism, compressed saves and dataset formats. | Orchestra-Research/ | 13k | 8 repos | ~1.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 36 | 36.Rlt Refactor Review and refactor inference-runtime code using concrete rules for responsibility boundaries, state ownership, interfaces, asynchronous lifetimes, KV management, and maintainability. | ThinkFlowLab/ | 147 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | today |
| 37 | Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout. | Orchestra-Research/ | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 38 | Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation. | guqiong96/ | 465 | 1 repo | ~831 | Automated safety check: Pass | Apache-2.0 | 18 days ago |
| 39 | Review and adapt vllmmetax/registry registrations, quantization configurations, CustomOps and kernel dispatch against a target vLLM revision and installed MetaX APIs. | MetaX-MACA/ | 180 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 40 | Advisory guidance for Amazon EKS architecture and configuration decisions — compute strategy, networking, security, reliability, cost, autoscaling, observability, multi-tenancy, and upgrade planning. | aws-samples/ | 115 | — | ~5k | Automated safety check: Pass | MIT-0 | 2 days ago |
| 41 | Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding. | marin-community/ | 3.9k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 42 | 42.Test Model Test LLM models served by xinfer for correctness, output quality, and performance. | guoqingbao/ | 334 | — | ~2.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 43 | Validate changed quantization math, packed formats, and inference kernels against a Torch oracle; adjudicate near-tie mismatches before rejecting optimized arithmetic. | ModelCloud/ | 1.3k | — | ~1.4k | Automated safety check: Pass | Unknown | today |
| 44 | Compose, edit, refactor, and validate Lego-RL train/eval/infer .env configs and reusable scripts/templates modules. | LegoX/ | 113 | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 45 | Interactively build, push or load, and deploy an airunway component (controller or any provider) to the cluster | ai-runway/ | 102 | — | ~927 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 46 | Sets up a structured debugging session for a Dynamo bug — pull the report from a Linear ticket, GitHub issue, or pasted text, capture the environment, create a persistent worklog markdown file, and… | ai-dynamo/ | 8.3k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 47 | Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. | vllm-project/ | 7.1k | — | ~8.7k | Automated safety check: Pass | Apache-2.0 | today |
| 48 | Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. | vllm-project/ | 102 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |