AI model or service
vLLM agent skills, page 2
vLLM skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Review PRs and local changes for hsliuustc0106/vllm-rlt: Ouro engine and KV correctness, serving behavior, and BF16 accuracy/speed evidence. | ThinkFlowLab/ | 140 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 50 | 50.Check Model Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer. | guoqingbao/ | 333 | — | ~3.8k | Automated safety check: Pass | MIT | 28 days ago |
| 51 | Quick-reference help for fine-tuning language models with Axolotl, covering YAML configs, FSDP, context parallelism, compressed saves and dataset formats. | Orchestra-Research/ | 13k | 9 repos | ~1.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 52 | Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout. | Orchestra-Research/ | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 53 | 53.Add Recipe Add or update an in-repository vLLM-Omni model recipe with verified task, input, output, hardware, command, feature, and validation contracts. | vllm-project/ | 7.1k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 54 | Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation. | guqiong96/ | 464 | 1 repo | ~831 | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 55 | Review and adapt vllmmetax/registry registrations, quantization configurations, CustomOps and kernel dispatch against a target vLLM revision and installed MetaX APIs. | MetaX-MACA/ | 179 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 56 | 56.Test Model Test LLM models served by xinfer for correctness, output quality, and performance. | guoqingbao/ | 333 | — | ~2.6k | Automated safety check: Pass | MIT | 28 days ago |
| 57 | Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding. | marin-community/ | 3.9k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 58 | Validate changed quantization math, packed formats, and inference kernels against a Torch oracle; adjudicate near-tie mismatches before rejecting optimized arithmetic. | ModelCloud/ | 1.3k | — | ~1.4k | Automated safety check: Pass | Unknown | today |
| 59 | Compose, edit, refactor, and validate Lego-RL train/eval/infer .env configs and reusable scripts/templates modules. | LegoX/ | 108 | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | today |
| 60 | Interactively build, push or load, and deploy an airunway component (controller or any provider) to the cluster | ai-runway/ | 101 | — | ~927 | Automated safety check: Pass | Apache-2.0 | 11 days ago |
| 61 | Sets up a structured debugging session for a Dynamo bug — pull the report from a Linear ticket, GitHub issue, or pasted text, capture the environment, create a persistent worklog markdown file, and… | ai-dynamo/ | 8.2k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 62 | Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. | vllm-project/ | 7.1k | — | ~8.7k | Automated safety check: Pass | Apache-2.0 | today |
| 63 | Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. | vllm-project/ | 103 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 64 | Produces ranked, evidence-grounded next steps when an ML experiment has stalled, drawing on a Leeroopedia knowledge base or on fetched docs and issues. | Leeroo-AI/ | 195 | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 65 | Build vLLM (CPU) end-to-end with the oneDNN + ZenDNN (zen64) backend via direct integration (NO zentorch). | amd/ | 158 | — | ~5.2k | Automated safety check: Pass | Unknown | yesterday |
| 66 | Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills. | amd/ | 398 | — | ~4k | Automated safety check: Notes | MIT | today |
| 67 | A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and… | NVIDIA-AI-Blueprints/ | 1.9k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 68 | Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. | Orchestra-Research/ | 13k | 3 repos | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 69 | Half-Quadratic Quantization for LLMs without calibration data. | Orchestra-Research/ | 13k | 3 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 70 | High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 3 repos | ~2.1k | Automated safety check: Notes | MIT | 3 mo ago |
| 71 | Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads. | Orchestra-Research/ | 13k | 3 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 72 | Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs. | Orchestra-Research/ | 13k | 3 repos | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 73 | Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends. | Orchestra-Research/ | 13k | 3 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 74 | Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor… | agentsope/ | 459 | — | ~3k | Automated safety check: Pass | MIT | 1 mo ago |
| 75 | Find evidence-backed simplification candidates in vLLM-Omni and, when requested, turn them into focused proposals or code changes. | vllm-project/ | 7.1k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | today |
| 76 | 76.Recif Eval Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines. | sohu-mptc/ | 107 | — | ~915 | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 77 | Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. | vllm-project/ | 103 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 78 | Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim. | Leeroo-AI/ | 195 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 79 | Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models. | amd/ | 181 | — | ~3k | Automated safety check: Pass | MIT | 10 days ago |
| 80 | verl-omni commit message + PR conventions and the mandatory contribution policy. | verl-project/ | 1.2k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 81 | 81.Dev Bump Bump Heron version via the VERSION-file SSOT. An agent skill from Netis/heron. | Netis/ | 102 | — | ~983 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 82 | Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs. | MetaX-MACA/ | 179 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 83 | This is an experimental Skill. An agent skill from GoogleCloudPlatform/accelerated-platforms. | GoogleCloudPlatform/ | 106 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 84 | Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding… | azrtydxb/ | 108 | — | ~980 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 85 | Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works. | vllm-project/ | 7.1k | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 86 | Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda. | amd/ | 398 | — | ~5.7k | Automated safety check: Notes | MIT | today |
| 87 | Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. | vllm-project/ | 103 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 88 | Sync documentation with project state using ICAV workflow. An agent skill from Netis/heron. | Netis/ | 102 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 89 | Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups. | Orchestra-Research/ | 13k | 3 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 90 | Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision. | MetaX-MACA/ | 179 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 91 | Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect. | BBuf/ | 911 | — | ~3.9k | Automated safety check: Pass | No licence | 2 days ago |
| 92 | Generate and run tests for vllm-project/vllm-omni with CI-aligned levels and markers; wire new tests into Buildkite (test-ready.yml for L1/L2, test-merge.yml for L3, test-nightly.yml for L4). | vllm-project/ | 7.1k | — | ~17k | Automated safety check: Pass | Apache-2.0 | today |
| 93 | Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API. | vllm-project/ | 103 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 94 | Optimize Ollama configuration for the current machine's hardware. | luongnv89/ | 131 | — | ~4.1k | Automated safety check: Notes | MIT | today |
| 95 | Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer. | amd/ | 398 | — | ~1.7k | Automated safety check: Notes | MIT | today |
| 96 | Establish the shared environment, source/runtime correspondence, target confirmation and validation evidence for MetaX vLLM upgrades. | MetaX-MACA/ | 179 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 8 days ago |