Search
Python · LLM inference and serving
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release. | R6410418/ | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 2 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | yesterday |
| 3 | 3.One Eval 驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。 | OpenDCAI/ | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 4 | Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable. | EGalahad/ | 146 | — | ~1.1k | Automated safety check: Pass | No licence | 12 days ago |
| 5 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | yesterday |
| 6 | Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models. | Orchestra-Research/ | 13k | 9 repos | ~4k | Automated safety check: Pass | MIT | 3 mo ago |
| 7 | Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout. | Orchestra-Research/ | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 8 | Naming conventions for SGLang speculative decoding identifiers. | sgl-project/ | 37k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 10 | Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling. | oracle/ | 125 | — | ~2.4k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 11 | GGUF format and llama.cpp quantization for efficient CPU/GPU inference. | Orchestra-Research/ | 13k | 3 repos | ~2.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 12 | Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK. | oracle/ | 125 | — | ~1.4k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 13 | Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model. | Orchestra-Research/ | 13k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 14 | Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 15 | Explains three ways to speed up LLM inference: draft-model speculative decoding, Medusa heads and lookahead decoding with Jacobi iteration, and when each one fits. | Orchestra-Research/ | 13k | 2 repos | ~3.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 16 | Explains RWKV, a hybrid that trains in parallel like a GPT and runs inference like an RNN with constant memory per token, plus usage, fine-tuning and troubleshooting. | Orchestra-Research/ | 13k | 2 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 17 | Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API. | vllm-project/ | 102 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 18 | Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer. | amd/ | 408 | — | ~1.7k | Automated safety check: Notes | MIT | today |
| 19 | Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA. | BBuf/ | 938 | — | ~7.5k | Automated safety check: Pass | No licence | 5 days ago |
| 20 | Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning. | amd/ | 182 | — | ~1.4k | Automated safety check: Pass | MIT | 12 days ago |
| 21 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 22 | 22.Llama Cpp llama.cpp local GGUF inference + HF Hub model discovery. An agent skill from Tommy-yw/RunbookHermes. | Tommy-yw/ | 546 | 4 repos | ~2.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 23 | Build private, on-device AI features on iPhone, iPad, and Mac with Foundation Models, Core ML, MLX Swift, or llama.cpp. | dpearson2699/ | 1.2k | — | ~3.4k | Automated safety check: Pass | Unknown | 2 mo ago |
| 24 | Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts. | amd/ | 182 | — | ~4.8k | Automated safety check: Pass | MIT | 12 days ago |
| 25 | Python backend for Databricks Apps — FastAPI (default), Flask, Dash, Streamlit, Gradio, Reflex. | databricks/ | 345 | — | ~3.1k | Automated safety check: Pass | Unknown | today |
| 26 | Run OpenMed models fully on-device with the MLX (Apple Silicon), CoreML (iOS/macOS), or ONNX/WebGPU (cross-platform/browser) backends, including convert-quantize-run workflows. | maziyarpanahi/ | 5.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 27 | Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation. | mirage-project/ | 2.5k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 28 | Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom. | AMD-AGI/ | 219 | — | ~7.1k | Automated safety check: Notes | Unknown | today |
| 29 | End-to-end ONNX PTQ workflow for AMD Quark — from a .onnx file (and calibration data) to a quantized .onnx output. | amd/ | 182 | — | ~4.5k | Automated safety check: Pass | MIT | 12 days ago |
| 30 | NVIDIA DeepStream SDK development with Python pyservicemaker API. | NVIDIA/ | 3.6k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 31 | Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. | amd/ | 182 | — | ~1.9k | Automated safety check: Notes | MIT | 12 days ago |
| 32 | Installs or verifies AMD Quark and ensures the selected Python environment has an accelerator-matched PyTorch. | amd/ | 182 | — | ~3.5k | Automated safety check: Notes | MIT | 12 days ago |
| 33 | 33.Setup Set up, install, and configure CONFIDE local de-identification — installs Python deps (natasha, scrubadub, phonenumbers, pymorphy2), ensures Ollama + pulls the default qwen2.5:3b model, detects… | glebis/ | 391 | — | ~1.1k | Automated safety check: Pass | MIT | 2 days ago |
| 34 | 34.Vllm Ascend vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | today |
| 35 | 35.Litellm Operate, configure, secure, and troubleshoot the LiteLLM AI gateway (proxy) and Python SDK: run the proxy (litellm --config), route to 100+ providers through one OpenAI-compatible API, configure… | magnus919/ | 115 | — | ~4.2k | Automated safety check: Notes | MIT | today |
| 36 | 36.Llama Cpp Run quantized LLMs locally with llama.cpp — CPU+GPU inference, GGUF format, OpenAI-compatible server, and Python bindings. | AlexAI-MCP/ | 135 | — | ~2.3k | Automated safety check: Pass | MIT | 6 mo ago |
| 37 | 昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持… | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | today |
| 38 | Fetch current model names from AI providers (Anthropic, OpenAI, Gemini, Ollama), classify them into tiers (fast/default/heavy), and detect new models. | aiskillstore/ | 433 | — | ~1.9k | Automated safety check: Pass | No licence | today |
| 39 | 39.Llama Cpp Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems. | magnus919/ | 115 | — | ~2.3k | Automated safety check: Pass | MIT | today |