Search
Docker · LLM inference and serving
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 2 | Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit. | vllm-project/ | 2.9k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 4 | 4.Deploy Deploy BankMCP™ to a small server so it works in claude.ai and on the phone: Railway or Fly.io, volume, domain, setup page, connector. | noskillish/ | 277 | — | ~744 | Automated safety check: Pass | MIT | 13 days ago |
| 5 | Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout. | Orchestra-Research/ | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 7 | Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills. | amd/ | 408 | — | ~4k | Automated safety check: Notes | MIT | 2 days ago |
| 8 | Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda. | amd/ | 408 | — | ~5.7k | Automated safety check: Notes | MIT | 2 days ago |
| 9 | Routes a NanoClaw agent group to a local Ollama model instead of the Anthropic API, using environment overrides and an optional block on Anthropic hosts. | nanocoai/ | 31k | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 10 | 10.Vss Deploy Deploys and manages VSS through setup.sh and its Docker Compose overlays. | open-edge-platform/ | 171 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 11 | Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI. | oracle/ | 125 | — | ~1.5k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 12 | Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server. | vllm-project/ | 102 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | 6 mo ago |
| 13 | 13.Vllm Server Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | 2 days ago |
| 14 | Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom. | AMD-AGI/ | 219 | — | ~7.1k | Automated safety check: Notes | Unknown | yesterday |
| 15 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.6k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 16 | Build Chat Question and Answer Core Docker images from source using direct Docker or Docker Compose build commands (backend CPU, backend GPU, backend Ollama, and UI). | open-edge-platform/ | 171 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 17 | Deploy Chat Question-and-Answer Core with Docker Compose (OpenVINO CPU, OpenVINO GPU, or Ollama CPU), including env setup, profile selection, startup verification, health checks, and teardown. | open-edge-platform/ | 171 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 18 | Troubleshoot Chat Question-and-Answer Core end-to-end across Docker Compose and Helm deployments, including startup failures, health/API errors, runtime mismatches (OpenVINO vs Ollama), model/config… | open-edge-platform/ | 171 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 19 | Extend, test, debug, or integrate the Model Download microservice codebase. | open-edge-platform/ | 171 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 20 | A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether… | NVIDIA/ | 3.6k | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 21 | The mandatory pre-launch gate and four-verb execution contract for every TAO workflow or action. | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 22 | Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline). | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 23 | End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks. | amd/ | 182 | — | ~6.2k | Automated safety check: Pass | MIT | 13 days ago |
| 24 | 24.Ollama Setup Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Notes | MIT | yesterday |
| 25 | 25.Vllm Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API… | magnus919/ | 115 | — | ~4.1k | Automated safety check: Notes | MIT | yesterday |
| 26 | Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |