Search
Kubernetes · LLM inference and serving
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 2 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 3 | Advisory guidance for Amazon EKS architecture and configuration decisions — compute strategy, networking, security, reliability, cost, autoscaling, observability, multi-tenancy, and upgrade planning. | aws-samples/ | 115 | — | ~5k | Automated safety check: Pass | MIT-0 | 3 days ago |
| 4 | Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. | vllm-project/ | 102 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 5 | Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | 2 days ago |
| 6 | Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. | sickn33/ | 47k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | 2 days ago |
| 7 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | sickn33/ | 47k | 1 repo | ~2.3k | Automated safety check: Pass | MIT | 2 days ago |
| 8 | Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters. | google/ | 21k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 9 | Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. | google/ | 21k | — | ~2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 10 | Set up AI Runway on AKS — from bare cluster to running model. | microsoft/ | 255 | 1 repo | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 11 | Debug live Azure Kubernetes Service (AKS) incidents with a read-only, evidence-first investigation. | microsoft/ | 255 | — | ~3.2k | Automated safety check: Pass | MIT | yesterday |
| 12 | The mandatory pre-launch gate and four-verb execution contract for every TAO workflow or action. | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 13 | 13.Vllm Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API… | magnus919/ | 115 | — | ~4.1k | Automated safety check: Notes | MIT | yesterday |
| 14 | Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. | BagelHole/ | 1.2k | — | ~2k | Automated safety check: Pass | MIT | 4 mo ago |
| 15 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |