Search
vLLM
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Produces ranked, evidence-grounded next steps when an ML experiment has stalled, drawing on a Leeroopedia knowledge base or on fetched docs and issues. | Leeroo-AI/ | 195 | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 50 | Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills. | amd/ | 408 | — | ~4k | Automated safety check: Notes | MIT | yesterday |
| 51 | Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling. | oracle/ | 125 | — | ~2.4k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 52 | Build vLLM (CPU) end-to-end with the oneDNN + ZenDNN (zen64) backend via direct integration (NO zentorch). | amd/ | 158 | — | ~5.2k | Automated safety check: Pass | Unknown | 4 days ago |
| 53 | Review PRs and local changes for hsliuustc0106/vllm-rlt: Ouro engine and KV correctness, serving behavior, and BF16 accuracy/speed evidence. | ThinkFlowLab/ | 149 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 54 | Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor… | agentsope/ | 436 | — | ~3k | Automated safety check: Pass | MIT | 2 days ago |
| 55 | Find evidence-backed simplification candidates in vLLM-Omni and, when requested, turn them into focused proposals or code changes. | vllm-project/ | 7.1k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | today |
| 56 | 56.Recif Eval Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines. | sohu-mptc/ | 107 | — | ~915 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 57 | Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. | vllm-project/ | 102 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 58 | Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim. | Leeroo-AI/ | 195 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 59 | Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models. | amd/ | 182 | — | ~3k | Automated safety check: Pass | MIT | 13 days ago |
| 60 | verl-omni commit message + PR conventions and the mandatory contribution policy. | verl-project/ | 1.2k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 61 | Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK. | oracle/ | 125 | — | ~1.4k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 62 | 62.Dev Bump Bump Heron version via the VERSION-file SSOT. An agent skill from Netis/heron. | Netis/ | 102 | — | ~983 | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 63 | A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and… | NVIDIA-AI-Blueprints/ | 1.9k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 64 | Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. | Orchestra-Research/ | 13k | 2 repos | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 65 | Half-Quadratic Quantization for LLMs without calibration data. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 66 | High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 2 repos | ~2.1k | Automated safety check: Notes | MIT | 3 mo ago |
| 67 | Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 68 | Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs. | Orchestra-Research/ | 13k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 69 | Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 70 | Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs. | MetaX-MACA/ | 180 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 71 | This is an experimental Skill. An agent skill from GoogleCloudPlatform/accelerated-platforms. | GoogleCloudPlatform/ | 106 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 72 | Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding… | azrtydxb/ | 108 | — | ~980 | Automated safety check: Pass | Apache-2.0 | today |
| 73 | Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works. | vllm-project/ | 7.1k | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 74 | Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda. | amd/ | 408 | — | ~5.7k | Automated safety check: Notes | MIT | yesterday |
| 75 | 75.Rlt Perf Opt Analyze and optimize vllm-rlt inference performance using reproducible unprofiled benchmarks, paired ops-only/full profiles, source-level attribution, and correctness checks. | ThinkFlowLab/ | 149 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | today |
| 76 | Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. | vllm-project/ | 102 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 77 | Sync documentation with project state using ICAV workflow. An agent skill from Netis/heron. | Netis/ | 102 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 78 | Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect. | BBuf/ | 938 | — | ~3.9k | Automated safety check: Pass | No licence | 6 days ago |
| 79 | Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision. | MetaX-MACA/ | 180 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 80 | Generate and run tests for vllm-project/vllm-omni with CI-aligned levels and markers; wire new tests into Buildkite (test-ready.yml for L1/L2, test-merge.yml for L3, test-nightly.yml for L4). | vllm-project/ | 7.1k | — | ~17k | Automated safety check: Pass | Apache-2.0 | today |
| 81 | Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API. | vllm-project/ | 102 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 82 | Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer. | amd/ | 408 | — | ~1.7k | Automated safety check: Notes | MIT | yesterday |
| 83 | Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups. | Orchestra-Research/ | 13k | 2 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 84 | Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA. | BBuf/ | 938 | — | ~7.5k | Automated safety check: Pass | No licence | 6 days ago |
| 85 | Establish the shared environment, source/runtime correspondence, target confirmation and validation evidence for MetaX vLLM upgrades. | MetaX-MACA/ | 180 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 86 | 86.Vss Deploy Deploys and manages VSS through setup.sh and its Docker Compose overlays. | open-edge-platform/ | 171 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 87 | This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. | vllm-project/ | 102 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 88 | Optimize Ollama configuration for the current machine's hardware. | luongnv89/ | 131 | — | ~4.1k | Automated safety check: Notes | MIT | today |
| 89 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 90 | A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence. | BBuf/ | 938 | — | ~1.5k | Automated safety check: Pass | No licence | 6 days ago |
| 91 | 91.Upgrade Deps Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL. | areal-project/ | 5.8k | — | ~6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 92 | A skill your agent uses whenever a developer needs to deploy VSS to Kubernetes, helm install VSS, configure values.yaml for VSS, or run VSS on k8s with GPU/vLLM for the… | open-edge-platform/ | 171 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 93 | Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI. | oracle/ | 125 | — | ~1.5k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 94 | DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity… | open-thoughts/ | 301 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 95 | Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval. | henryalouf/ | 157 | — | ~1.6k | Automated safety check: Pass | MIT | 4 mo ago |
| 96 | Trim large MetaX model directories for dummy smoke tests or real-checkpoint loading on limited GPUs. | MetaX-MACA/ | 180 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |