Search
AI & LLM Engineering · vLLM
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Select and verify the current region-specific serving container URI for a SageMaker model deployment. | waybarrios/ | 534 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 98 | Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL… | open-thoughts/ | 301 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 99 | Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware. | sickn33/ | 47k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 100 | Curated upstream guidance for Huggingface Community Evals; use when the workflow matches the user goal. | sickn33/ | 47k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 101 | Build, test, and debug Hermes Agent RL environments for Atropos training. | Tommy-yw/ | 546 | — | ~3.3k | Automated safety check: Pass | MIT | 4 mo ago |
| 102 | 102.Local LLM Expert Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. | sickn33/ | 47k | 2 repos | ~1.6k | Automated safety check: Pass | MIT | 2 days ago |
| 103 | 103.Vllm Server Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | 2 days ago |
| 104 | Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure. | NVIDIA-AI-Blueprints/ | 1.9k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 105 | Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation. | mirage-project/ | 2.5k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 106 | 106.Hyperloom Setup Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom. | AMD-AGI/ | 219 | — | ~7.1k | Automated safety check: Notes | Unknown | yesterday |
| 107 | 107.Vllm Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads. | Prism-Shadow/ | 2.5k | — | ~1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 108 | Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 109 | Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard. | Xilinx/ | 150 | — | ~3.4k | Automated safety check: Pass | MIT | yesterday |
| 110 | 110.Model Serving LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components. | ancoleman/ | 525 | — | ~3.4k | Automated safety check: Pass | MIT | 10 mo ago |
| 111 | Root-cause a vLLM torch-nightly CI regression report. An agent skill from pytorch/test-infra. | pytorch/ | 113 | — | ~2.4k | Automated safety check: Pass | Unknown | yesterday |
| 112 | Add and manage evaluation results in Hugging Face model cards. | sickn33/ | 47k | 2 repos | ~418 | Automated safety check: Pass | MIT | 2 days ago |
| 113 | Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data. | NVIDIA/ | 3.6k | 1 repo | ~2.3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 114 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.6k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 115 | 115.LLM Gateway Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking. | sickn33/ | 47k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | 2 days ago |
| 116 | Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. | sickn33/ | 47k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | 2 days ago |
| 117 | vLLM: high-throughput LLM serving, OpenAI API, quantization. | Luciole-Studio/ | 171 | 2 repos | ~2.3k | Automated safety check: Pass | MIT | 3 days ago |
| 118 | lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 2 repos | ~3.1k | Automated safety check: Pass | MIT | 3 days ago |
| 119 | Neutral, framework-agnostic decision tree for project kickoff: "which agent / RAG / LLM framework should I reach for?" Synthesizes the ecosystem sections of 7 landmark-project SOPs (LangGraph… | agentsope/ | 436 | — | ~5.8k | Automated safety check: Pass | MIT | 2 days ago |
| 120 | Review vLLM-routed torch-nightly root causes against existing upstream vLLM issues using read-only search results, and emit a validated upstream-checks artifact for the filer. | pytorch/ | 113 | — | ~1.3k | Automated safety check: Pass | Unknown | yesterday |
| 121 | Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson. | NVIDIA/ | 3.6k | 1 repo | ~2.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 122 | Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output. | NVIDIA/ | 3.6k | 1 repo | ~3.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 123 | 123.Datagen Launch Launch a datagen (trace-generation) job on an HPC cluster (Jupiter/Leonardo/Perlmutter) via the OpenThoughts-Agent hpc.launch --jobtype datagen entrypoint — the cluster-AGNOSTIC general flow… | open-thoughts/ | 301 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 124 | Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN… | grafana/ | 282 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 125 | 125.Cost Local Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a… | ruvnet/ | 74k | — | ~336 | Automated safety check: Notes | MIT | yesterday |
| 126 | Set up AI Runway on AKS — from bare cluster to running model. | microsoft/ | 255 | 1 repo | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 127 | [omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the… | rlaope/ | 3.2k | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 128 | Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin. | NVIDIA/ | 3.6k | 1 repo | ~3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 129 | Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck. | NVIDIA/ | 3.6k | 1 repo | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 130 | Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes. | open-thoughts/ | 301 | — | ~947 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 131 | Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy… | open-thoughts/ | 301 | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 132 | A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether… | NVIDIA/ | 3.6k | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 133 | Cross-engine decision rubric for self-hosting or recommending an LLM serving stack. | agentsope/ | 436 | — | ~6.1k | Automated safety check: Pass | MIT | 2 days ago |
| 134 | Project-kickoff rubric for the self-host vs managed-cloud decision — when is running your own inference engine / LLM platform worth the ops cost vs paying per-token for a managed API? | agentsope/ | 436 | — | ~6.3k | Automated safety check: Pass | MIT | 2 days ago |
| 135 | 135.Agentsop Vllm Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy. | agentsope/ | 436 | — | ~6.1k | Automated safety check: Pass | MIT | 2 days ago |
| 136 | 136.Engine Vllm Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it… | autonomous-ai/ | 1.2k | — | ~1.9k | Automated safety check: Pass | MIT | yesterday |
| 137 | Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent. | open-thoughts/ | 301 | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 138 | EXECUTE a staged codebase plan (from code-create-staged-plan or an existing notes/<codebase/ plan) one stage at a time, gate-by-gate, while keeping the local clone ground truth and a dated… | open-thoughts/ | 301 | — | ~977 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 139 | Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics. | benchflow-ai/ | 1.8k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 140 | How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks. | NVIDIA/ | 3.6k | — | ~5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 141 | End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks. | amd/ | 182 | — | ~6.2k | Automated safety check: Pass | MIT | 13 days ago |
| 142 | This skill should be used when the user asks about "local models", "custom models", "fine-tuning", "self-hosting models", "model selection", "which model should I use", "data privacy and models"… | Habitat-Thinking/ | 114 | — | ~1k | Automated safety check: Pass | Unknown | 21 days ago |
| 143 | 143.Vllm Ascend vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |
| 144 | Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode… | ascend-ai-coding/ | 174 | — | ~731 | Automated safety check: Pass | No licence | yesterday |