Search
AI & LLM Engineering · Kubernetes · For devops and sre engineers
- AI & LLM Engineering (remove filter)
- Kubernetes (remove filter)
- For devops and sre engineers (remove filter)
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 2 | Runs OpenAI Codex CLI as a non-interactive worker for CI, Docker, Kubernetes or remote servers, with sandbox modes and JSONL-friendly output. | XiaomiMiMo/ | 14k | — | ~2.7k | Automated safety check: Pass | MIT | 2 days ago |
| 3 | Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost. | Orchestra-Research/ | 13k | 4 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 4 | 4.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | 2 days ago |
| 5 | Creates, lists, updates and deletes Volcano Jobs and Queues in KubeSphere, with YAML templates for PyTorch, TensorFlow, MPI and batch jobs plus scheduling troubleshooting. | kubesphere/ | 17k | — | ~5.6k | Automated safety check: Pass | Unknown | 2 mo ago |
| 6 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | sickn33/ | 47k | 1 repo | ~2.3k | Automated safety check: Pass | MIT | 2 days ago |
| 7 | Deploy, monitor, and debug long GPU jobs on RENTED/remote instances (AutoDL, RunPod, vast.ai, Lambda, Slurm, K8s): teardown/billing safety, spot resilience, resumable checkpointing, OOM/NaN triage. | sickn33/ | 47k | 1 repo | ~5.8k | Automated safety check: Pass | MIT | 2 days ago |
| 8 | Debug live Azure Kubernetes Service (AKS) incidents with a read-only, evidence-first investigation. | microsoft/ | 255 | — | ~3.2k | Automated safety check: Pass | MIT | yesterday |
| 9 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |
| 10 | Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar. | NVIDIA/ | 3.6k | — | ~3.8k | Automated safety check: Warn | Apache-2.0 | 2 days ago |