Search

Docker · LLM inference and serving

26 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

dstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.02 days ago
2

Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

vllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0today
3

Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.

dstackai/dstack2.3k—~403Automated safety check: PassMPL-2.02 days ago
4

Deploy BankMCP™ to a small server so it works in claude.ai and on the phone: Railway or Fly.io, volume, domain, setup page, connector.

noskillish/bankmcp277—~744Automated safety check: PassMIT13 days ago
5

Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

Orchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT3 mo ago
6

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

Orchestra-Research/AI-Research-SKILLs13k4 repos~1.3kAutomated safety check: PassMIT3 mo ago
7

Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

amd/skills408—~4kAutomated safety check: NotesMIT2 days ago
8

Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.

amd/skills408—~5.7kAutomated safety check: NotesMIT2 days ago
9

Routes a NanoClaw agent group to a local Ollama model instead of the Anthropic API, using environment overrides and an optional block on Anthropic hosts.

nanocoai/nanoclaw31k—~1.5kAutomated safety check: PassMITyesterday
10

Deploys and manages VSS through setup.sh and its Docker Compose overlays.

open-edge-platform/edge-ai-libraries171—~4.1kAutomated safety check: PassApache-2.0yesterday
11
11.Aqua MetricsOfficial

Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI.

oracle/accelerated-data-science125—~1.5kAutomated safety check: PassUPL-1.01 mo ago
12

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

vllm-project/vllm-skills102—~2.5kAutomated safety check: NotesApache-2.06 mo ago
13

Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

sickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT2 days ago
14

Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.

AMD-AGI/Hyperloom219—~7.1kAutomated safety check: NotesUnknownyesterday
15
15.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.6k1 repo~1.8kAutomated safety check: PassApache-2.02 days ago
16

Build Chat Question and Answer Core Docker images from source using direct Docker or Docker Compose build commands (backend CPU, backend GPU, backend Ollama, and UI).

open-edge-platform/edge-ai-libraries171—~1.2kAutomated safety check: PassApache-2.0yesterday
17

Deploy Chat Question-and-Answer Core with Docker Compose (OpenVINO CPU, OpenVINO GPU, or Ollama CPU), including env setup, profile selection, startup verification, health checks, and teardown.

open-edge-platform/edge-ai-libraries171—~2.9kAutomated safety check: PassApache-2.0yesterday
18

Troubleshoot Chat Question-and-Answer Core end-to-end across Docker Compose and Helm deployments, including startup failures, health/API errors, runtime mismatches (OpenVINO vs Ollama), model/config…

open-edge-platform/edge-ai-libraries171—~2.6kAutomated safety check: PassApache-2.0yesterday
19

Extend, test, debug, or integrate the Model Download microservice codebase.

open-edge-platform/edge-ai-libraries171—~2.8kAutomated safety check: PassApache-2.0yesterday
20
20.Deepstream SopOfficial

A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether…

NVIDIA/skills3.6k—~4.7kAutomated safety check: NotesApache-2.02 days ago
21

The mandatory pre-launch gate and four-verb execution contract for every TAO workflow or action.

NVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.02 days ago
22

Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).

NVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.02 days ago
23

End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks.

amd/Quark182—~6.2kAutomated safety check: PassMIT13 days ago
24

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: NotesMITyesterday
25
25.Vllm

Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

magnus919/agent-skills115—~4.1kAutomated safety check: NotesMITyesterday
26

Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licenceyesterday