Search

LLM inference and serving

373 skills found, page 5.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
193

当你需要配置/新增/排查 LLM 提供商(OpenAI/Azure/Gemini/Anthropic/Ollama/各类兼容网关)时使用;确保 YAML/Schema/代码一致。

xrkseek/XRK-AGT138—~304Automated safety check: PassMIT10 days ago
194

Integrates local AI capabilities into applications using Embeddable Lemonade.

amd/skills408—~6kAutomated safety check: PassMIT2 days ago
195

A skill your agent uses when testing or benchmarking target/draft GGUF pairs for speculative decoding compatibility, tokenizer agreement, draft acceptance rate, or staged verification behavior.

Mesh-LLM/mesh-llm3.5k—~260Automated safety check: PassApache-2.0today
196

Guidance for implementing batching schedulers for LLM inference systems with compilation-based accelerators.

lazyFrogLOL/Harness_Engineering128—~2.2kAutomated safety check: PassNo licence4 mo ago
197

Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio.

sickn33/agentic-awesome-skills47k2 repos~1.6kAutomated safety check: PassMIT2 days ago
198

Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

sickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT2 days ago
199

Extract structured domain knowledge from AI models in-session or from local open-source models via Ollama.

sickn33/agentic-awesome-skills47k2 repos~926Automated safety check: PassMIT2 days ago
200

Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure.

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.5kAutomated safety check: PassApache-2.0yesterday
201

Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation.

mirage-project/mirage2.5k—~4.6kAutomated safety check: PassApache-2.03 days ago
202

Measure local AI task latency, token usage, errors and verified outcomes using Pudu AI hardware evidence and installed Ollama models.

davila7/claude-code-templates33k—~1.3kAutomated safety check: PassMITyesterday
203

Reduce the size of glTF/GLB 3D assets to cut Git LFS bandwidth/storage while keeping them loadable by the editor's Bevy 0.18 glTF loader.

elodin-sys/elodin547—~1kAutomated safety check: PassApache-2.0yesterday
204

Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.

AMD-AGI/Hyperloom219—~7.1kAutomated safety check: NotesUnknownyesterday
205
205.Ollama

Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.

Prism-Shadow/penguin-harness2.5k—~839Automated safety check: NotesApache-2.0yesterday
206
206.Vllm

Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.

Prism-Shadow/penguin-harness2.5k—~1kAutomated safety check: PassApache-2.0yesterday
207

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
208

Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard.

Xilinx/mlir-air150—~3.4kAutomated safety check: PassMITyesterday
209

LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components.

ancoleman/ai-design-components525—~3.4kAutomated safety check: PassMIT10 mo ago
210

Diagnoses Qdrant search quality issues. An agent skill from github/awesome-copilot.

github/awesome-copilot40k1 repo~928Automated safety check: PassMIT2 days ago
211

Optimize AgentDB performance with quantization (4-32x memory reduction), HNSW indexing (150x faster search), caching, and batch operations.

aiskillstore/marketplace4336 repos~3kAutomated safety check: PassNo licenceyesterday
212

Root-cause a vLLM torch-nightly CI regression report. An agent skill from pytorch/test-infra.

pytorch/test-infra113—~2.4kAutomated safety check: PassUnknownyesterday
213

Add and manage evaluation results in Hugging Face model cards.

sickn33/agentic-awesome-skills47k2 repos~418Automated safety check: PassMIT2 days ago
214
214.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.6k1 repo~1.8kAutomated safety check: PassApache-2.02 days ago
215

Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

sickn33/agentic-awesome-skills47k1 repo~2.5kAutomated safety check: PassMIT2 days ago
216

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMIT2 days ago
217

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

sickn33/agentic-awesome-skills47k1 repo~2.3kAutomated safety check: PassMIT2 days ago
218

vLLM: high-throughput LLM serving, OpenAI API, quantization.

Luciole-Studio/Misaka-Agent1712 repos~2.3kAutomated safety check: PassMIT3 days ago
219

Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters.

google/skills21k—~3.1kAutomated safety check: PassApache-2.0yesterday
220

Diagnoses and improves Qdrant search relevance. An agent skill from github/awesome-copilot.

github/awesome-copilot40k1 repo~336Automated safety check: PassMIT2 days ago
221

Neutral, framework-agnostic decision tree for project kickoff: "which agent / RAG / LLM framework should I reach for?" Synthesizes the ecosystem sections of 7 landmark-project SOPs (LangGraph…

agentsope/SkillAlchemy436—~5.8kAutomated safety check: PassMIT2 days ago
222

Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
223

Review vLLM-routed torch-nightly root causes against existing upstream vLLM issues using read-only search results, and emit a validated upstream-checks artifact for the filer.

pytorch/test-infra113—~1.3kAutomated safety check: PassUnknownyesterday
224

Command-line interface for Ollama - Local LLM inference and model management via Ollama REST API.

HKUDS/CLI-Anything52k—~1.2kAutomated safety check: PassApache-2.019 days ago
225

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

NVIDIA/skills3.6k1 repo~2.9kAutomated safety check: PassApache-2.02 days ago
226

Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

NVIDIA/skills3.6k1 repo~3.1kAutomated safety check: PassApache-2.02 days ago
227

Google platform decision and setup guidance, loaded on demand from Google's skill catalog.

google/skills21k—~1.7kAutomated safety check: PassApache-2.0yesterday
228
228.Gke InferenceOfficial

Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.

google/skills21k—~2kAutomated safety check: PassApache-2.0yesterday
229

Inspect a target ONNX model and prepare metadata for Quark ONNX PTQ planning.

amd/Quark182—~4.3kAutomated safety check: PassMIT13 days ago
230

Launch a datagen (trace-generation) job on an HPC cluster (Jupiter/Leonardo/Perlmutter) via the OpenThoughts-Agent hpc.launch --jobtype datagen entrypoint — the cluster-AGNOSTIC general flow…

open-thoughts/OpenThoughts-Agent301—~2.5kAutomated safety check: PassApache-2.013 days ago
231

TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama.

secondsky/claude-skills227—~3.6kAutomated safety check: NotesMIT13 days ago
232
232.ML AIOfficial

Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN…

grafana/skills282—~1.3kAutomated safety check: PassApache-2.03 days ago
233

Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a…

ruvnet/ruflo74k—~336Automated safety check: NotesMITyesterday
234

Validate ChatQnA Core REST APIs from docs/user-guide/api-reference.md using repeatable curl-based smoke tests, runtime-specific endpoint checks (OpenVINO or Ollama), and concise pass/fail evidence.

open-edge-platform/edge-ai-libraries171—~1.6kAutomated safety check: PassApache-2.0yesterday
235

Build Chat Question and Answer Core Docker images from source using direct Docker or Docker Compose build commands (backend CPU, backend GPU, backend Ollama, and UI).

open-edge-platform/edge-ai-libraries171—~1.2kAutomated safety check: PassApache-2.0yesterday
236

Deploy Chat Question-and-Answer Core with Docker Compose (OpenVINO CPU, OpenVINO GPU, or Ollama CPU), including env setup, profile selection, startup verification, health checks, and teardown.

open-edge-platform/edge-ai-libraries171—~2.9kAutomated safety check: PassApache-2.0yesterday
237

Troubleshoot Chat Question-and-Answer Core end-to-end across Docker Compose and Helm deployments, including startup failures, health/API errors, runtime mismatches (OpenVINO vs Ollama), model/config…

open-edge-platform/edge-ai-libraries171—~2.6kAutomated safety check: PassApache-2.0yesterday
238

High-throughput LLM inference on NVIDIA GPUs. An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1711 repo~1.3kAutomated safety check: PassMIT3 days ago
239

Search PubMed, preserve complete abstracts and sections, and produce research briefings with abstract openings or optional OpenAI/Ollama summaries.

ClawBio/ClawBio1.2k—~3.3kAutomated safety check: PassMIT2 days ago
240

EDM and club music as a style layer (电子舞曲风格层) - a form built on layer addition and subtraction rather than verses and choruses.

jtydhr88/music-composition-skills154—~2.1kAutomated safety check: PassMIT19 days ago