Agent skill

Vllm Deploy K8s

by vllm-project in vllm-project/vllm-skills

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Vllm Deploy K8s

skills CLI
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vllm-project/vllm-skills vllm-deploy-k8s --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-k8s .claude/skills/vllm-deploy-k8s && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-deploy-k8s
GitHub stars
102
Token cost
~2k tokens
SKILL.md length
764 words
Files
3
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

  • Works in 5 steps: Check HF token secret → Check if deployment already exists → Deploy → …
  • The user wants to deploy
  • SKILL.md covers What this skill does, Prerequisites, Deployment Steps and Default Configuration, plus 5 more sections
  • Calls kubectl, curl and python3; needs HF_TOKEN

What it does

Vllm Deploy K8s is an agent skill from vllm-project/vllm-skills. Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Use this skill whenever the user wants to deploy, run, or serve vLLM on a Kubernetes cluster, including creating deployments, services, checking existing deployments, or managing vLLM on K8s.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `templates/vllm-deployment.yaml` and `templates/vllm-service.yaml`).

It sits in AI & LLM Engineering, covering Container orchestration, LLM inference and serving and Deployment. It works with Kubernetes, vLLM, OpenAI and Hugging Face. The repository describes itself as: Agent skills for vLLM. The licence is Apache-2.0.

When your agent uses it

  • The user wants to deploy
  • Serve vLLM on a Kubernetes cluster
  • Including creating deployments
  • Checking existing deployments

Example prompts

  • “/vllm-deploy-k8s”

Requirements

  • Python 3
  • Docker

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Check HF token secret
  2. Check if deployment already exists
  3. Deploy
  4. Wait and verify
  5. Print deployment summary

What it can do on your machine

Read from SKILL.md and the folder at commit c996234. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • curl
    • python3
    • hf

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.vllm.ai
    • hub.docker.com
    • docs.nvidia.com
    • kubernetes.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm Deploy K8s loads about 2k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 764 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vllm-project/vllm-skills at commit c996234, republished under its Apache-2.0 licence (© vllm-project). 764 words, ~1,966 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-deploy-k8s/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
vllm-deploy-k8s
description
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Use this skill whenever the user wants to deploy, run, or serve vLLM on a Kubernetes cluster, including creating deployments, services, checking existing deployments, or managing vLLM on K8s.

vLLM Kubernetes Deployment

A Claude skill for deploying vLLM to Kubernetes using YAML templates. Deploys a vLLM OpenAI-compatible server as a Kubernetes Deployment with a ClusterIP Service, GPU resources, and health probes.

What this skill does

  • Deploy vLLM as a Kubernetes Deployment + Service with NVIDIA GPU support
  • Check if a vLLM deployment already exists before deploying
  • Check if the Hugging Face token secret exists, and ask the user for their token if not
  • Use the vllm/vllm-openai:latest image by default (user can specify a different version)
  • Provide sensible default configuration that users can customize (model, replicas, GPU count, extra vLLM flags, etc.)

Prerequisites

  • kubectl configured with access to a Kubernetes cluster
  • NVIDIA GPU Operator or device plugin installed on cluster nodes
  • Hugging Face token (required for gated models like Llama, optional for public models)

Deployment Steps

Step 1: Check HF token secret

Before deploying, check if the hf-token Kubernetes secret exists in the target namespace:

bash
kubectl get secret hf-token -n <namespace>
  • If the secret exists: proceed to Step 2.
  • If the secret does not exist: ask the user to provide their Hugging Face token, then create the secret:
bash
kubectl create secret generic hf-token --from-literal=HF_TOKEN="<user-provided-token>" -n <namespace>

This is required for gated models (e.g., meta-llama/Meta-Llama-3.1-8B). For public models, the secret is optional but recommended to avoid rate limits.

Step 2: Check if deployment already exists

Before applying, check if a vLLM deployment already exists:

bash
kubectl get deployment vllm -n <namespace>
  • If it exists: inform the user that the deployment already exists. Show the current image and status. Ask the user if they want to update it or skip.
  • If it does not exist: proceed to deploy.
Step 3: Deploy

Apply the template YAML files to deploy vLLM:

bash
kubectl apply -f templates/vllm-service.yaml -n <namespace>
kubectl apply -f templates/vllm-deployment.yaml -n <namespace>
Step 4: Wait and verify

Wait for the deployment to roll out:

bash
kubectl rollout status deployment/vllm -n <namespace> --timeout=600s

Verify the pod is running and ready:

bash
kubectl get pods -n <namespace> -l app=vllm

Confirm the pod shows READY 1/1 and STATUS Running. If the pod is not ready yet, wait and check again. If it's in CrashLoopBackOff or Error, check the logs with kubectl logs -n <namespace> -l app=vllm.

Step 5: Print deployment summary

Once the pod is ready, print a summary message to the user in this format (replace placeholders with actual values):

🎉 **vLLM Deployment Successful!**

| Resource | Name | Status |
|----------|------|--------|
| Deployment | <deployment-name> | <ready>/<total> Ready |
| Service | <service-name> | ClusterIP:<port> |
| Pod | <pod-name> | Running |
| Image | <image> | |
| Model | <model> | |

&nbsp;

**To test the API, run these two commands in your terminal:**

**1. Open a port-forward** (this connects your local port <port> to the vLLM service inside the cluster):

kubectl port-forward svc/vllm-svc <port>:<port> -n <namespace>

**2. In a separate terminal**, send a test request to the OpenAI-compatible API:

curl -s http://localhost:<port>/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"<model>","messages":[{"role":"user","content":"Hello!"}],"max_tokens":50}' | python3 -m json.tool

If everything is working, you'll get a JSON response with the model's reply.

Default Configuration

The templates use the following defaults:

ParameterDefault Value
Imagevllm/vllm-openai:latest
ModelQwen/Qwen2.5-1.5B-Instruct
Port8000
Replicas1
GPU count1
GPU memory utilization0.85
Tensor parallel size1
CPU request / limit12 / 128
Memory request / limit100Gi / 400Gi
Shared memory (dshm)80Gi
Show full SKILL.md (375 more words)Show less

Customization

When the user requests changes, modify the template YAML files before applying. The following can be customized:

  • Image version: Change image: vllm/vllm-openai:<version> in templates/vllm-deployment.yaml (default: latest). Use a specific version tag like v0.17.1 if the user requests it.
  • Model: Change the model name in the vllm serve command inside the Deployment args.
  • Extra vLLM flags: Append additional flags to the vllm serve command in the Deployment args (e.g., --max-model-len 4096, --kv-cache-dtype fp8, --enforce-eager, --generation-config vllm).
  • Replicas: Change replicas: in the Deployment spec.
  • GPU count: Change nvidia.com/gpu in both requests and limits under resources.
  • Tensor parallel size: Change --tensor-parallel-size flag to match the GPU count.
  • CPU/Memory resources: Change cpu and memory values under requests and limits.
  • Port: Change containerPort in the Deployment, port/targetPort in the Service, the port in all health probes (liveness, readiness, startup), AND add --port <port> to the vllm serve command in args. All four must match.
  • Namespace: Apply to a specific namespace using -n <namespace>.
  • Shared memory size: Change the sizeLimit of the dshm emptyDir volume.

Edit the template files using the Edit tool, then apply the modified templates.

Status Check

bash
kubectl get deployment,svc,pods -n <namespace> -l app=vllm

Cleanup

When the user asks to clean up or delete the vLLM deployment, run the following steps:

  1. Delete the Deployment and Service:
bash
kubectl delete -f templates/vllm-deployment.yaml -n <namespace>
kubectl delete -f templates/vllm-service.yaml -n <namespace>
  1. Ask the user if they also want to delete the HF token secret. If yes:
bash
kubectl delete secret hf-token -n <namespace>
  1. Verify everything is cleaned up:
bash
kubectl get deployment,svc,pods -n <namespace> -l app=vllm
  1. Print a summary message to the user:
vLLM deployment has been cleaned up from namespace <namespace>.
Deleted: Deployment/vllm, Service/vllm-svc
HF token secret: <kept/deleted>

Troubleshooting

  • Pod stuck in Pending: No GPU nodes available. Check kubectl describe pod <pod-name> for scheduling errors. Ensure NVIDIA GPU Operator or device plugin is installed.
  • Pod OOMKilled: Increase memory limits in the Deployment, or use a smaller model.
  • ImagePullBackOff: Check the image name and tag. Verify the node has access to Docker Hub / the container registry.
  • Startup probe failures (CrashLoopBackOff): Model download may be slow. Check logs with kubectl logs <pod-name>. Ensure hf-token secret exists for gated models. Increase failureThreshold on the startup probe if needed.
  • HF_TOKEN not working: Verify the secret exists: kubectl get secret hf-token -n <namespace>. Check the token is valid.
  • GPU not detected in container: Ensure nvidia.com/gpu resource is requested and the NVIDIA device plugin is running on the node.

References

© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in plugins/vllm-skills/skills/vllm-deploy-k8s of vllm-project/vllm-skills.

  • SKILL.md
  • templates/vllm-deployment.yaml
  • templates/vllm-service.yaml

Open the folder on GitHubat commit c996234

Compare with similar skills

Vllm Deploy K8s next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Deploy K8s compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Deploy K8s this skillvllm-project/vllm-skills102—~2kAutomated safety check: PassApache-2.0
Gke Manifest Generationgoogle/skills21k—~3.1kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT
LLM Inference Scalingsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
LLM Inference ScalingBagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT

Similar skills

  • Official

    Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters.

    21k GitHub stars~3.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 5 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Inference Scaling

    sickn33/agentic-awesome-skills

    Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Inference Scaling

    BagelHole/DevOps-Security-Agent-Skills

    Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

    1.2k GitHub stars~2k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed

More from vllm-project/vllm-skills

  • Vllm Bench Random Synthetic

    vllm-project/vllm-skills

    Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

    102 GitHub stars~1.5k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Bench Serve

    vllm-project/vllm-skills

    Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

    102 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy Simple

    vllm-project/vllm-skills

    Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

    102 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Prefix Cache Bench

    vllm-project/vllm-skills

    This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.

    102 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    102 GitHub stars~2.5k tokensUpdated 6 mo ago
    Auto-check: notes

Questions about Vllm Deploy K8s

What does Vllm Deploy K8s do?

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Vllm Deploy K8s is an agent skill from vllm-project/vllm-skills. Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

When should I use Vllm Deploy K8s?

Vllm Deploy K8s fits situations like: the user wants to deploy; serve vLLM on a Kubernetes cluster; including creating deployments; checking existing deployments.

How do I install Vllm Deploy K8s in Claude Code?

Run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a claude-code`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-deploy-k8s in vllm-project/vllm-skills) into .claude/skills/vllm-deploy-k8s in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Deploy K8s in Codex?

Run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a codex`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-deploy-k8s in vllm-project/vllm-skills) into .agents/skills/vllm-deploy-k8s in your project. Codex loads it when a task matches its description.

Can I use Vllm Deploy K8s in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-k8s -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-deploy-k8s, .gemini/skills/vllm-deploy-k8s, .github/skills/vllm-deploy-k8s and .opencode/skills/vllm-deploy-k8s in your project.

What does Vllm Deploy K8s need to run?

Going by SKILL.md and its folder, Vllm Deploy K8s needs the command-line tools its instructions call (kubectl, curl, python3 and hf) and credentials named HF_TOKEN. Our summary lists: Python 3; Docker.

Does Vllm Deploy K8s access the network?

SKILL.md names 4 domains. As links in the text: docs.vllm.ai, hub.docker.com, docs.nvidia.com and kubernetes.io. This is read from the text; nothing was executed.

Is Vllm Deploy K8s safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm Deploy K8s use?

Vllm Deploy K8s is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm Deploy K8s use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vllm Deploy K8s?

Skills that share tags, products or a category with Vllm Deploy K8s: Gke Manifest Generation (google/skills, 21k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and LLM Inference Scaling (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Deploy K8s?

vllm-project (a GitHub organization) maintains it in vllm-project/vllm-skills, which has 102 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 3, 2026.

Source: vllm-project/vllm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.