Search

Python · LLM inference and serving

39 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

R6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT3 mo ago
2

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

dstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0yesterday
3

驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

OpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.01 mo ago
4

Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable.

EGalahad/sim2real146—~1.1kAutomated safety check: PassNo licence12 days ago
5

Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.

dstackai/dstack2.3k—~403Automated safety check: PassMPL-2.0yesterday
6

Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models.

Orchestra-Research/AI-Research-SKILLs13k9 repos~4kAutomated safety check: PassMIT3 mo ago
7

Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

Orchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT3 mo ago
8

Naming conventions for SGLang speculative decoding identifiers.

sgl-project/sglang37k2 repos~1.6kAutomated safety check: PassApache-2.0today
9

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

Orchestra-Research/AI-Research-SKILLs13k4 repos~1.3kAutomated safety check: PassMIT3 mo ago
10
10.Aqua DeploymentOfficial

Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.

oracle/accelerated-data-science125—~2.4kAutomated safety check: PassUPL-1.01 mo ago
11

GGUF format and llama.cpp quantization for efficient CPU/GPU inference.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.6kAutomated safety check: PassMIT3 mo ago
12

Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

oracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.01 mo ago
13

Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.5kAutomated safety check: PassMIT3 mo ago
14

Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
15

Explains three ways to speed up LLM inference: draft-model speculative decoding, Medusa heads and lookahead decoding with Jacobi iteration, and when each one fits.

Orchestra-Research/AI-Research-SKILLs13k2 repos~3.5kAutomated safety check: PassMIT3 mo ago
16

Explains RWKV, a hybrid that trains in parallel like a GPT and runs inference like an RNN with constant memory per token, plus usage, fine-tuning and troubleshooting.

Orchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT3 mo ago
17

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

vllm-project/vllm-skills102—~1.6kAutomated safety check: PassApache-2.06 mo ago
18

Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.

amd/skills408—~1.7kAutomated safety check: NotesMITtoday
19

Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~7.5kAutomated safety check: PassNo licence5 days ago
20

Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

amd/Quark182—~1.4kAutomated safety check: PassMIT12 days ago
21

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills408—~2.3kAutomated safety check: PassMITtoday
22

llama.cpp local GGUF inference + HF Hub model discovery. An agent skill from Tommy-yw/RunbookHermes.

Tommy-yw/RunbookHermes5464 repos~2.2kAutomated safety check: PassMIT4 mo ago
23

Build private, on-device AI features on iPhone, iPad, and Mac with Foundation Models, Core ML, MLX Swift, or llama.cpp.

dpearson2699/swift-ios-skills1.2k—~3.4kAutomated safety check: PassUnknown2 mo ago
24

Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts.

amd/Quark182—~4.8kAutomated safety check: PassMIT12 days ago
25

Python backend for Databricks Apps — FastAPI (default), Flask, Dash, Streamlit, Gradio, Reflex.

databricks/databricks-agent-skills345—~3.1kAutomated safety check: PassUnknowntoday
26

Run OpenMed models fully on-device with the MLX (Apple Silicon), CoreML (iOS/macOS), or ONNX/WebGPU (cross-platform/browser) backends, including convert-quantize-run workflows.

maziyarpanahi/openmed5.5k—~2kAutomated safety check: PassApache-2.0today
27

Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation.

mirage-project/mirage2.5k—~4.6kAutomated safety check: PassApache-2.02 days ago
28

Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.

AMD-AGI/Hyperloom219—~7.1kAutomated safety check: NotesUnknowntoday
29

End-to-end ONNX PTQ workflow for AMD Quark — from a .onnx file (and calibration data) to a quantized .onnx output.

amd/Quark182—~4.5kAutomated safety check: PassMIT12 days ago
30
30.Deepstream DevOfficial

NVIDIA DeepStream SDK development with Python pyservicemaker API.

NVIDIA/skills3.6k—~3.3kAutomated safety check: PassApache-2.0yesterday
31

Diagnose failed Quark installation, PTQ execution, script generation, or export attempts.

amd/Quark182—~1.9kAutomated safety check: NotesMIT12 days ago
32

Installs or verifies AMD Quark and ensures the selected Python environment has an accelerator-matched PyTorch.

amd/Quark182—~3.5kAutomated safety check: NotesMIT12 days ago
33

Set up, install, and configure CONFIDE local de-identification — installs Python deps (natasha, scrubadub, phonenumbers, pymorphy2), ensures Ollama + pulls the default qwen2.5:3b model, detects…

glebis/claude-skills391—~1.1kAutomated safety check: PassMIT2 days ago
34

vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licencetoday
35

Operate, configure, secure, and troubleshoot the LiteLLM AI gateway (proxy) and Python SDK: run the proxy (litellm --config), route to 100+ providers through one OpenAI-compatible API, configure…

magnus919/agent-skills115—~4.2kAutomated safety check: NotesMITtoday
36

Run quantized LLMs locally with llama.cpp — CPU+GPU inference, GGUF format, OpenAI-compatible server, and Python bindings.

AlexAI-MCP/hermes-CCC135—~2.3kAutomated safety check: PassMIT6 mo ago
37

昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持…

ascend-ai-coding/awesome-ascend-skills174—~1.2kAutomated safety check: PassNo licencetoday
38

Fetch current model names from AI providers (Anthropic, OpenAI, Gemini, Ollama), classify them into tiers (fast/default/heavy), and detect new models.

aiskillstore/marketplace433—~1.9kAutomated safety check: PassNo licencetoday
39

Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems.

magnus919/agent-skills115—~2.3kAutomated safety check: PassMITtoday