Gke Manifest Generation
google/skills
Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters.
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-k8s --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-k8s .claude/skills/vllm-deploy-k8s && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vllm-deploy-k8s" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-k8s into .claude/skills/vllm-deploy-k8s/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-k8s", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-k8sType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-k8s --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-k8s .agents/skills/vllm-deploy-k8s && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vllm-deploy-k8s" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-k8s into .agents/skills/vllm-deploy-k8s/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-k8s", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-k8s --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-k8s .cursor/skills/vllm-deploy-k8s && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vllm-deploy-k8s" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-k8s into .cursor/skills/vllm-deploy-k8s/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-k8s", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vllm-project/vllm-skills.git --path plugins/vllm-skills/skills/vllm-deploy-k8s--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-k8s --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-k8s .gemini/skills/vllm-deploy-k8s && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vllm-deploy-k8s" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-k8s into .gemini/skills/vllm-deploy-k8s/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-k8s", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vllm-project/vllm-skills vllm-deploy-k8sInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-k8s .github/skills/vllm-deploy-k8s && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vllm-deploy-k8s" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-k8s into .github/skills/vllm-deploy-k8s/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-k8s", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-k8s --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-k8s .opencode/skills/vllm-deploy-k8s && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vllm-deploy-k8s" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-k8s into .opencode/skills/vllm-deploy-k8s/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-k8s", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vllm-deploy-k8sDeploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
Vllm Deploy K8s is an agent skill from vllm-project/vllm-skills. Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Use this skill whenever the user wants to deploy, run, or serve vLLM on a Kubernetes cluster, including creating deployments, services, checking existing deployments, or managing vLLM on K8s.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `templates/vllm-deployment.yaml` and `templates/vllm-service.yaml`).
It sits in AI & LLM Engineering, covering Container orchestration, LLM inference and serving and Deployment. It works with Kubernetes, vLLM, OpenAI and Hugging Face. The repository describes itself as: Agent skills for vLLM. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c996234. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlcurlpython3hfFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.vllm.aihub.docker.comdocs.nvidia.comkubernetes.ioFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vllm Deploy K8s loads about 2k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 764 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vllm-project/vllm-skills at commit c996234, republished under its Apache-2.0 licence (© vllm-project). 764 words, ~1,966 tokens.
.claude/skills/vllm-deploy-k8s/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.A Claude skill for deploying vLLM to Kubernetes using YAML templates. Deploys a vLLM OpenAI-compatible server as a Kubernetes Deployment with a ClusterIP Service, GPU resources, and health probes.
vllm/vllm-openai:latest image by default (user can specify a different version)kubectl configured with access to a Kubernetes clusterBefore deploying, check if the hf-token Kubernetes secret exists in the target namespace:
kubectl get secret hf-token -n <namespace>kubectl create secret generic hf-token --from-literal=HF_TOKEN="<user-provided-token>" -n <namespace>This is required for gated models (e.g., meta-llama/Meta-Llama-3.1-8B). For public models, the secret is optional but recommended to avoid rate limits.
Before applying, check if a vLLM deployment already exists:
kubectl get deployment vllm -n <namespace>Apply the template YAML files to deploy vLLM:
kubectl apply -f templates/vllm-service.yaml -n <namespace>
kubectl apply -f templates/vllm-deployment.yaml -n <namespace>Wait for the deployment to roll out:
kubectl rollout status deployment/vllm -n <namespace> --timeout=600sVerify the pod is running and ready:
kubectl get pods -n <namespace> -l app=vllmConfirm the pod shows READY 1/1 and STATUS Running. If the pod is not ready yet, wait and check again. If it's in CrashLoopBackOff or Error, check the logs with kubectl logs -n <namespace> -l app=vllm.
Once the pod is ready, print a summary message to the user in this format (replace placeholders with actual values):
🎉 **vLLM Deployment Successful!**
| Resource | Name | Status |
|----------|------|--------|
| Deployment | <deployment-name> | <ready>/<total> Ready |
| Service | <service-name> | ClusterIP:<port> |
| Pod | <pod-name> | Running |
| Image | <image> | |
| Model | <model> | |
**To test the API, run these two commands in your terminal:**
**1. Open a port-forward** (this connects your local port <port> to the vLLM service inside the cluster):
kubectl port-forward svc/vllm-svc <port>:<port> -n <namespace>
**2. In a separate terminal**, send a test request to the OpenAI-compatible API:
curl -s http://localhost:<port>/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"<model>","messages":[{"role":"user","content":"Hello!"}],"max_tokens":50}' | python3 -m json.tool
If everything is working, you'll get a JSON response with the model's reply.The templates use the following defaults:
| Parameter | Default Value |
|---|---|
| Image | vllm/vllm-openai:latest |
| Model | Qwen/Qwen2.5-1.5B-Instruct |
| Port | 8000 |
| Replicas | 1 |
| GPU count | 1 |
| GPU memory utilization | 0.85 |
| Tensor parallel size | 1 |
| CPU request / limit | 12 / 128 |
| Memory request / limit | 100Gi / 400Gi |
| Shared memory (dshm) | 80Gi |
When the user requests changes, modify the template YAML files before applying. The following can be customized:
image: vllm/vllm-openai:<version> in templates/vllm-deployment.yaml (default: latest). Use a specific version tag like v0.17.1 if the user requests it.vllm serve command inside the Deployment args.vllm serve command in the Deployment args (e.g., --max-model-len 4096, --kv-cache-dtype fp8, --enforce-eager, --generation-config vllm).replicas: in the Deployment spec.nvidia.com/gpu in both requests and limits under resources.--tensor-parallel-size flag to match the GPU count.cpu and memory values under requests and limits.containerPort in the Deployment, port/targetPort in the Service, the port in all health probes (liveness, readiness, startup), AND add --port <port> to the vllm serve command in args. All four must match.-n <namespace>.sizeLimit of the dshm emptyDir volume.Edit the template files using the Edit tool, then apply the modified templates.
kubectl get deployment,svc,pods -n <namespace> -l app=vllmWhen the user asks to clean up or delete the vLLM deployment, run the following steps:
kubectl delete -f templates/vllm-deployment.yaml -n <namespace>
kubectl delete -f templates/vllm-service.yaml -n <namespace>kubectl delete secret hf-token -n <namespace>kubectl get deployment,svc,pods -n <namespace> -l app=vllmvLLM deployment has been cleaned up from namespace <namespace>.
Deleted: Deployment/vllm, Service/vllm-svc
HF token secret: <kept/deleted>kubectl describe pod <pod-name> for scheduling errors. Ensure NVIDIA GPU Operator or device plugin is installed.memory limits in the Deployment, or use a smaller model.kubectl logs <pod-name>. Ensure hf-token secret exists for gated models. Increase failureThreshold on the startup probe if needed.kubectl get secret hf-token -n <namespace>. Check the token is valid.nvidia.com/gpu resource is requested and the NVIDIA device plugin is running on the node.© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in plugins/vllm-skills/skills/vllm-deploy-k8s of vllm-project/vllm-skills.
Open the folder on GitHubat commit c996234
Vllm Deploy K8s next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vllm Deploy K8s this skillvllm-project/vllm-skills | 102 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Gke Manifest Generationgoogle/skills | 21k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| vLLM Model ServingOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | |
| LLM Inference Scalingsickn33/agentic-awesome-skills | 47k | 1 repos | ~2.1k | Automated safety check: Pass | MIT | |
| LLM Inference ScalingBagelHole/DevOps-Security-Agent-Skills | 1.2k | — | ~2k | Automated safety check: Pass | MIT |
google/skills
Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
sickn33/agentic-awesome-skills
Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.
BagelHole/DevOps-Security-Agent-Skills
Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.
huggingface/skills
Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.
vllm-project/vllm-skills
Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.
vllm-project/vllm-skills
Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.
vllm-project/vllm-skills
Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.
vllm-project/vllm-skills
This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
Works with
Categories
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Vllm Deploy K8s is an agent skill from vllm-project/vllm-skills. Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
Vllm Deploy K8s fits situations like: the user wants to deploy; serve vLLM on a Kubernetes cluster; including creating deployments; checking existing deployments.
Run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a claude-code`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-deploy-k8s in vllm-project/vllm-skills) into .claude/skills/vllm-deploy-k8s in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a codex`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-deploy-k8s in vllm-project/vllm-skills) into .agents/skills/vllm-deploy-k8s in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-deploy-k8s, .gemini/skills/vllm-deploy-k8s, .github/skills/vllm-deploy-k8s and .opencode/skills/vllm-deploy-k8s in your project.
Going by SKILL.md and its folder, Vllm Deploy K8s needs the command-line tools its instructions call (kubectl, curl, python3 and hf) and credentials named HF_TOKEN. Our summary lists: Python 3; Docker.
SKILL.md names 4 domains. As links in the text: docs.vllm.ai, hub.docker.com, docs.nvidia.com and kubernetes.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Vllm Deploy K8s is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Vllm Deploy K8s: Gke Manifest Generation (google/skills, 21k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and LLM Inference Scaling (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vllm-project (a GitHub organization) maintains it in vllm-project/vllm-skills, which has 102 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 3, 2026.
Source: vllm-project/vllm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.