Topic · AI & LLM Engineering
Best LLM inference and serving skills, page 7
LLM inference and serving skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 289 | Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent. | amd/ | 181 | — | ~2.3k | Automated safety check: Pass | MIT | 10 days ago |
| 290 | Validate Quark quantization output using four lightweight checks: auxiliary file copy alignment, excluded tensor MD5 byte-identity, config.json deep comparison after stripping quantization keys, and… | amd/ | 181 | — | ~1.6k | Automated safety check: Pass | MIT | 10 days ago |
| 291 | Route Quark user goals to the correct atomic skill or workflow. | amd/ | 181 | — | ~1.9k | Automated safety check: Pass | MIT | 10 days ago |
| 292 | Route Quark user goals to the correct atomic skill or workflow. | amd/ | 181 | — | ~1.8k | Automated safety check: Pass | MIT | 10 days ago |
| 293 | Detect upstream Quark changes that affect the skill system and classify required updates. | amd/ | 181 | — | ~1.6k | Automated safety check: Pass | MIT | 10 days ago |
| 294 | Extract structured domain knowledge from AI models in-session or from local open-source models via Ollama. | majiayu000/ | 666 | 3 repos | ~847 | Automated safety check: Pass | MIT | today |
| 295 | 在本机 Mac 或 Apple Silicon 上部署 Gemma 4 12B。本地安装/升级 llama.cpp,下载 GGUF 量化模型,用 llama-server 暴露 OpenAI-compatible API,或用 Ollama 暴露本地模型服务;按用户需求在默认 Q4KM、64K/128K 长上下文、QAT Q40 @ 256K、左右对比演示之间选择,配置 tmux… | majiayu000/ | 286 | — | ~875 | Automated safety check: Notes | MIT | today |
| 296 | 296.Performance Performance optimization patterns covering Core Web Vitals, React render optimization, lazy loading, image optimization, backend profiling, LLM inference, and sustainability UX. | yonatangross/ | 289 | — | ~3.5k | Automated safety check: Pass | MIT | today |
| 297 | 297.Train Fasttext This skill provides guidance for training FastText text classification models with constraints on accuracy and model size. | lazyFrogLOL/ | 128 | — | ~1.9k | Automated safety check: Pass | No licence | 4 mo ago |
| 298 | Diagnoses Qdrant search quality issues. An agent skill from qdrant/skills. | qdrant/ | 253 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 299 | Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or… | neo4j-contrib/ | 114 | — | ~5.6k | Automated safety check: Notes | MIT | yesterday |
| 300 | 300.Run Model Eval A skill your agent uses when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the… | heypinchy/ | 182 | — | ~2k | Automated safety check: Pass | AGPL-3.0 | 17 days ago |
| 301 | EXECUTE a staged codebase plan (from code-create-staged-plan or an existing notes/<codebase/ plan) one stage at a time, gate-by-gate, while keeping the local clone ground truth and a dated… | open-thoughts/ | 301 | — | ~977 | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 302 | Plan or review LoRA and edit-training work specifically for FLUX.2 Klein or Qwen-Image-Edit, including paired datasets, trainer-version contracts, and held-out fidelity checks. | AnastasiyaW/ | 154 | — | ~4.5k | Automated safety check: Pass | MIT | today |
| 303 | Selects, deploys, and customizes AI models on Amazon SageMaker. | aws/ | 2.8k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 304 | 304.Quark Install Installs or verifies AMD Quark and ensures the selected Python environment has an accelerator-matched PyTorch. | amd/ | 181 | — | ~3.5k | Automated safety check: Notes | MIT | 10 days ago |
| 305 | 305.Quark Torch Ptq Runs an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reproducible artifacts, request… | amd/ | 181 | — | ~2.3k | Automated safety check: Pass | MIT | 10 days ago |
| 306 | 306.Release Debrief Summarises the delta between a tool's latest release and the last summary the user saw. | sammcj/ | 162 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 307 | Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics. | benchflow-ai/ | 1.8k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 308 | How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks. | NVIDIA/ | 3.5k | — | ~5k | Automated safety check: Notes | Apache-2.0 | today |
| 309 | Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline). | NVIDIA/ | 3.5k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | today |
| 310 | Three-layer PII anonymization for session transcripts (therapy, coaching, consulting, mentoring). | glebis/ | 389 | — | ~841 | Automated safety check: Pass | MIT | 11 days ago |
| 311 | 311.Setup Set up, install, and configure CONFIDE local de-identification — installs Python deps (natasha, scrubadub, phonenumbers, pymorphy2), ensures Ollama + pulls the default qwen2.5:3b model, detects… | glebis/ | 389 | — | ~1.1k | Automated safety check: Pass | MIT | 11 days ago |
| 312 | 312.Forkmind A skill your agent uses when debugging, comparing, or regression-testing LLM / agent calls — when the user wants to capture LLM traffic, see a conversation as a branchable DAG, fork an alternative… | ccplugins/ | 968 | — | ~868 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 313 | 313.Ollama Setup Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Notes | MIT | today |
| 314 | End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks. | amd/ | 181 | — | ~6.2k | Automated safety check: Pass | MIT | 10 days ago |
| 315 | 315.Model Researcher Add new/custom AI models to opencode.json. An agent skill from IgorWarzocha/Opencode-Workflows. | IgorWarzocha/ | 122 | — | ~2.2k | Automated safety check: Pass | No licence | 8 mo ago |
| 316 | Add and manage evaluation results in Hugging Face model cards. | majiayu000/ | 666 | 3 repos | ~5.6k | Automated safety check: Notes | MIT | today |
| 317 | Trending Hugging Face models, datasets, and spaces — filtered by license sanity, dedup vs same-week quantizations, with a "why notable" line per pick (architecture shift, size step, license change… | BankrBot/ | 1.2k | — | ~496 | Automated safety check: Pass | No licence | 2 days ago |
| 318 | This skill should be used when the user asks about "local models", "custom models", "fine-tuning", "self-hosting models", "model selection", "which model should I use", "data privacy and models"… | Habitat-Thinking/ | 114 | — | ~1k | Automated safety check: Pass | Unknown | 17 days ago |
| 319 | 319.Pp Ollama Cloud Routes every prompt to the right hosted Ollama model. An agent skill from mvanhorn/printing-press-library. | mvanhorn/ | 2.1k | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | today |
| 320 | Rhythm section arranging and MIDI programming (节奏组与打ち込み) - drums, bass and harmony instruments as one unit, and how to make programmed parts sound played. | jtydhr88/ | 150 | — | ~3.2k | Automated safety check: Pass | MIT | 15 days ago |
| 321 | 321.Vllm Ascend vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | today |
| 322 | Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode… | ascend-ai-coding/ | 174 | — | ~731 | Automated safety check: Pass | No licence | today |
| 323 | Expert knowledge for Azure Databricks development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration, integrations &… | MicrosoftDocs/ | 777 | 1 repo | ~14k | Automated safety check: Pass | CC-BY-4.0 | 2 days ago |
| 324 | 324.Ollama A skill your agent uses when running open-weight LLMs locally with Ollama — pulling and tagging models, calling the local API, picking a quantization or GGUF, writing Modelfiles, and sizing VRAM and… | ericrisco/ | 167 | — | ~2.8k | Automated safety check: Pass | MIT | today |
| 325 | 325.RAG A skill your agent uses when building grounded Q&A over your own corpus — chunk, retrieve hybrid, rerank, ground, cite chunk ids, refuse when the sources fall short — or when the right document is… | ericrisco/ | 167 | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 326 | A skill your agent uses when calling open-weight LLMs on Together AI or Fireworks AI's OpenAI-compatible endpoints — baseurl plus namespaced model id, the cheapest model that clears the bar… | ericrisco/ | 167 | — | ~3.3k | Automated safety check: Pass | MIT | today |
| 327 | 327.Vector DB A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric… | ericrisco/ | 167 | — | ~2.8k | Automated safety check: Pass | MIT | today |
| 328 | 328.Vllm A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or… | majiayu000/ | 666 | 1 repo | ~3.6k | Automated safety check: Pass | MIT | today |
| 329 | 329.Vllm vLLM is a high-throughput inference and serving engine for large language models that exposes an OpenAI-compatible HTTP API and a Python batch API. | majiayu000/ | 666 | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 330 | 330.Zed Zed is a fast, GPU-accelerated code editor written in Rust, with built-in AI agent and edit predictions, real-time collaboration, Vim mode and WebAssembly extensions. | majiayu000/ | 666 | 1 repo | ~2.4k | Automated safety check: Notes | Apache-2.0 | today |
| 331 | 331.Litellm Operate, configure, secure, and troubleshoot the LiteLLM AI gateway (proxy) and Python SDK: run the proxy (litellm --config), route to 100+ providers through one OpenAI-compatible API, configure… | magnus919/ | 113 | — | ~4.2k | Automated safety check: Notes | MIT | yesterday |
| 332 | 332.Vllm Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API… | magnus919/ | 113 | — | ~4.1k | Automated safety check: Notes | MIT | yesterday |
| 333 | Summarizes arbitrarily long text (1k-1M words) using recursive map-reduce with any LLM backend. | swyxio/ | 175 | — | ~6.3k | Automated safety check: Pass | MIT | 3 days ago |
| 334 | 优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供… | majiayu000/ | 286 | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 335 | 335.Vllm Bench Serve Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve. | ascend-ai-coding/ | 174 | — | ~5.6k | Automated safety check: Pass | No licence | today |
| 336 | 336.AI Model Router Intelligent AI model router that automatically switches between two configured models (local for simple tasks, cloud for complex ones). | LeoYeAI/ | 2.2k | — | ~811 | Automated safety check: Pass | MIT | 2 mo ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- LLM evaluation308
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- GPU and accelerator computing176
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23