Search

AI & LLM Engineering · Kubernetes

26 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

dstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.02 days ago
2

Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.

dstackai/dstack2.3k—~403Automated safety check: PassMPL-2.02 days ago
3

Use FieldFlow to inspect and reduce noisy JSON CLI output before it reaches model context.

guillaumegay13/fieldflow110—~872Automated safety check: PassMIT2 mo ago
4

Runs OpenAI Codex CLI as a non-interactive worker for CI, Docker, Kubernetes or remote servers, with sandbox modes and JSONL-friendly output.

XiaomiMiMo/MiMo-Code14k—~2.7kAutomated safety check: PassMIT2 days ago
5

Compose, edit, refactor, and validate Lego-RL train/eval/infer .env configs and reusable scripts/templates modules.

LegoX/Lego-RL113—~2.1kAutomated safety check: NotesApache-2.03 days ago
6

Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost.

Orchestra-Research/AI-Research-SKILLs13k4 repos~2.4kAutomated safety check: PassMIT3 mo ago
7

dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.

dstackai/dstack2.3k—~6.2kAutomated safety check: WarnMPL-2.02 days ago
8

Use Langfuse's disposable per-PR previews at pr-N.preview.langfuse.com (synthetic data only).

langfuse/langfuse36k—~2.8kAutomated safety check: NotesUnknowntoday
9

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

vllm-project/vllm-skills102—~2kAutomated safety check: PassApache-2.06 mo ago
10

Creates, lists, updates and deletes Volcano Jobs and Queues in KubeSphere, with YAML templates for PyTorch, TensorFlow, MPI and batch jobs plus scheduling troubleshooting.

kubesphere/kubesphere17k—~5.6kAutomated safety check: PassUnknown2 mo ago
11

Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

sickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT2 days ago
12

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMIT2 days ago
13

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

sickn33/agentic-awesome-skills47k1 repo~2.3kAutomated safety check: PassMIT2 days ago
14

Deploy, monitor, and debug long GPU jobs on RENTED/remote instances (AutoDL, RunPod, vast.ai, Lambda, Slurm, K8s): teardown/billing safety, spot resilience, resumable checkpointing, OOM/NaN triage.

sickn33/agentic-awesome-skills47k1 repo~5.8kAutomated safety check: PassMIT2 days ago
15
15.Gke InferenceOfficial

Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.

google/skills21k—~2kAutomated safety check: PassApache-2.0yesterday
16

Discovers requirements, and generates architectural, design, and deployment guidance for a retrieval-augmented generation (RAG)-capable enterprise search system in Google Cloud.

google/skills21k—~4kAutomated safety check: PassApache-2.0yesterday
17

Set up AI Runway on AKS — from bare cluster to running model.

microsoft/GitHub-Copilot-for-Azure2551 repo~1.1kAutomated safety check: PassMITyesterday
18

Debug live Azure Kubernetes Service (AKS) incidents with a read-only, evidence-first investigation.

microsoft/GitHub-Copilot-for-Azure255—~3.2kAutomated safety check: PassMITyesterday
19

Manage GPU compute jobs on the Qizhi (启智) platform using qzcli — a kubectl-style CLI tool.

AI4Scientist/nano-scientist1282 repos~1.9kAutomated safety check: NotesNo licence4 mo ago
20

Safety guardrails for destructive commands. An agent skill from mr-daedalium/ostack-saas.

mr-daedalium/ostack-saas1141 repo~545Automated safety check: NotesMIT6 mo ago
21
21.Vllm

Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

magnus919/agent-skills115—~4.1kAutomated safety check: NotesMITyesterday
22

Diagnose and fix CoreWeave GPU scheduling, pod, and networking errors.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: PassMITyesterday
23

Deploy a GPU workload on CoreWeave with kubectl. An agent skill from jeremylongshore/tons-of-skills-marketplace.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: PassMITyesterday
24

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT4 mo ago
25

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2.1kAutomated safety check: PassMIT4 mo ago
26

Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar.

NVIDIA/skills3.6k—~3.8kAutomated safety check: WarnApache-2.0yesterday