Topic · AI & LLM Engineering

Best LLM inference and serving skills, page 7

Skills #289–336 of 372, ranked by score.

LLM inference and serving skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM inference and serving skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
289

Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent.

amd/Quark181—~2.3kAutomated safety check: PassMIT10 days ago
290

Validate Quark quantization output using four lightweight checks: auxiliary file copy alignment, excluded tensor MD5 byte-identity, config.json deep comparison after stripping quantization keys, and…

amd/Quark181—~1.6kAutomated safety check: PassMIT10 days ago
291

Route Quark user goals to the correct atomic skill or workflow.

amd/Quark181—~1.9kAutomated safety check: PassMIT10 days ago
292

Route Quark user goals to the correct atomic skill or workflow.

amd/Quark181—~1.8kAutomated safety check: PassMIT10 days ago
293

Detect upstream Quark changes that affect the skill system and classify required updates.

amd/Quark181—~1.6kAutomated safety check: PassMIT10 days ago
294

Extract structured domain knowledge from AI models in-session or from local open-source models via Ollama.

majiayu000/claude-skill-registry6663 repos~847Automated safety check: PassMITtoday
295

在本机 Mac 或 Apple Silicon 上部署 Gemma 4 12B。本地安装/升级 llama.cpp,下载 GGUF 量化模型,用 llama-server 暴露 OpenAI-compatible API,或用 Ollama 暴露本地模型服务;按用户需求在默认 Q4KM、64K/128K 长上下文、QAT Q40 @ 256K、左右对比演示之间选择,配置 tmux…

majiayu000/spellbook286—~875Automated safety check: NotesMITtoday
296

Performance optimization patterns covering Core Web Vitals, React render optimization, lazy loading, image optimization, backend profiling, LLM inference, and sustainability UX.

yonatangross/orchestkit289—~3.5kAutomated safety check: PassMITtoday
297

This skill provides guidance for training FastText text classification models with constraints on accuracy and model size.

lazyFrogLOL/Harness_Engineering128—~1.9kAutomated safety check: PassNo licence4 mo ago
298

Diagnoses Qdrant search quality issues. An agent skill from qdrant/skills.

qdrant/skills253—~2.3kAutomated safety check: PassApache-2.0today
299

Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or…

neo4j-contrib/neo4j-skills114—~5.6kAutomated safety check: NotesMITyesterday
300

A skill your agent uses when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the…

heypinchy/pinchy182—~2kAutomated safety check: PassAGPL-3.017 days ago
301

EXECUTE a staged codebase plan (from code-create-staged-plan or an existing notes/<codebase/ plan) one stage at a time, gate-by-gate, while keeping the local clone ground truth and a dated…

open-thoughts/OpenThoughts-Agent301—~977Automated safety check: PassApache-2.09 days ago
302

Plan or review LoRA and edit-training work specifically for FLUX.2 Klein or Qwen-Image-Edit, including paired datasets, trainer-version contracts, and held-out fidelity checks.

AnastasiyaW/codex-claude-code-config154—~4.5kAutomated safety check: PassMITtoday
303
303.AWS AI MLOfficial

Selects, deploys, and customizes AI models on Amazon SageMaker.

aws/agent-toolkit-for-aws2.8k—~1.7kAutomated safety check: PassApache-2.0today
304

Installs or verifies AMD Quark and ensures the selected Python environment has an accelerator-matched PyTorch.

amd/Quark181—~3.5kAutomated safety check: NotesMIT10 days ago
305

Runs an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reproducible artifacts, request…

amd/Quark181—~2.3kAutomated safety check: PassMIT10 days ago
306

Summarises the delta between a tool's latest release and the last summary the user saw.

sammcj/agentic-coding162—~1.8kAutomated safety check: PassApache-2.0today
307

Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics.

benchflow-ai/skillsbench1.8k—~2.3kAutomated safety check: PassApache-2.02 mo ago
308

How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.

NVIDIA/skills3.5k—~5kAutomated safety check: NotesApache-2.0today
309

Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).

NVIDIA/skills3.5k—~4.5kAutomated safety check: NotesApache-2.0today
310

Three-layer PII anonymization for session transcripts (therapy, coaching, consulting, mentoring).

glebis/claude-skills389—~841Automated safety check: PassMIT11 days ago
311
311.Setup

Set up, install, and configure CONFIDE local de-identification — installs Python deps (natasha, scrubadub, phonenumbers, pymorphy2), ensures Ollama + pulls the default qwen2.5:3b model, detects…

glebis/claude-skills389—~1.1kAutomated safety check: PassMIT11 days ago
312

A skill your agent uses when debugging, comparing, or regression-testing LLM / agent calls — when the user wants to capture LLM traffic, see a conversation as a branchable DAG, fork an alternative…

ccplugins/awesome-claude-code-plugins968—~868Automated safety check: PassApache-2.01 mo ago
313

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: NotesMITtoday
314

End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks.

amd/Quark181—~6.2kAutomated safety check: PassMIT10 days ago
315

Add new/custom AI models to opencode.json. An agent skill from IgorWarzocha/Opencode-Workflows.

IgorWarzocha/Opencode-Workflows122—~2.2kAutomated safety check: PassNo licence8 mo ago
316

Add and manage evaluation results in Hugging Face model cards.

majiayu000/claude-skill-registry6663 repos~5.6kAutomated safety check: NotesMITtoday
317

Trending Hugging Face models, datasets, and spaces — filtered by license sanity, dedup vs same-week quantizations, with a "why notable" line per pick (architecture shift, size step, license change…

BankrBot/skills1.2k—~496Automated safety check: PassNo licence2 days ago
318

This skill should be used when the user asks about "local models", "custom models", "fine-tuning", "self-hosting models", "model selection", "which model should I use", "data privacy and models"…

Habitat-Thinking/ai-literacy-superpowers114—~1kAutomated safety check: PassUnknown17 days ago
319

Routes every prompt to the right hosted Ollama model. An agent skill from mvanhorn/printing-press-library.

mvanhorn/printing-press-library2.1k—~2.9kAutomated safety check: NotesApache-2.0today
320

Rhythm section arranging and MIDI programming (节奏组与打ち込み) - drums, bass and harmony instruments as one unit, and how to make programmed parts sound played.

jtydhr88/music-composition-skills150—~3.2kAutomated safety check: PassMIT15 days ago
321

vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licencetoday
322

Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode…

ascend-ai-coding/awesome-ascend-skills174—~731Automated safety check: PassNo licencetoday
323
323.Azure DatabricksOfficial

Expert knowledge for Azure Databricks development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration, integrations &…

MicrosoftDocs/Agent-Skills7771 repo~14kAutomated safety check: PassCC-BY-4.02 days ago
324
324.Ollama

A skill your agent uses when running open-weight LLMs locally with Ollama — pulling and tagging models, calling the local API, picking a quantization or GGUF, writing Modelfiles, and sizing VRAM and…

ericrisco/rsc-harness167—~2.8kAutomated safety check: PassMITtoday
325
325.RAG

A skill your agent uses when building grounded Q&A over your own corpus — chunk, retrieve hybrid, rerank, ground, cite chunk ids, refuse when the sources fall short — or when the right document is…

ericrisco/rsc-harness167—~2.9kAutomated safety check: PassMITtoday
326

A skill your agent uses when calling open-weight LLMs on Together AI or Fireworks AI's OpenAI-compatible endpoints — baseurl plus namespaced model id, the cheapest model that clears the bar…

ericrisco/rsc-harness167—~3.3kAutomated safety check: PassMITtoday
327

A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…

ericrisco/rsc-harness167—~2.8kAutomated safety check: PassMITtoday
328
328.Vllm

A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…

majiayu000/claude-skill-registry6661 repo~3.6kAutomated safety check: PassMITtoday
329
329.Vllm

vLLM is a high-throughput inference and serving engine for large language models that exposes an OpenAI-compatible HTTP API and a Python batch API.

majiayu000/claude-skill-registry6661 repo~1.9kAutomated safety check: PassApache-2.0today
330
330.Zed

Zed is a fast, GPU-accelerated code editor written in Rust, with built-in AI agent and edit predictions, real-time collaboration, Vim mode and WebAssembly extensions.

majiayu000/claude-skill-registry6661 repo~2.4kAutomated safety check: NotesApache-2.0today
331

Operate, configure, secure, and troubleshoot the LiteLLM AI gateway (proxy) and Python SDK: run the proxy (litellm --config), route to 100+ providers through one OpenAI-compatible API, configure…

magnus919/agent-skills113—~4.2kAutomated safety check: NotesMITyesterday
332
332.Vllm

Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

magnus919/agent-skills113—~4.1kAutomated safety check: NotesMITyesterday
333

Summarizes arbitrarily long text (1k-1M words) using recursive map-reduce with any LLM backend.

swyxio/skills175—~6.3kAutomated safety check: PassMIT3 days ago
334

优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…

majiayu000/spellbook286—~1.1kAutomated safety check: PassMITtoday
335

Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve.

ascend-ai-coding/awesome-ascend-skills174—~5.6kAutomated safety check: PassNo licencetoday
336

Intelligent AI model router that automatically switches between two configured models (local for simple tasks, cloud for complex ones).

LeoYeAI/openclaw-master-skills2.2k—~811Automated safety check: PassMIT2 mo ago