Agent skill

Model Serving Kubernetes

by sickn33 in sickn33/agentic-awesome-skills

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

MITAuto-check passedDevOps & Cloud

Install Model Serving Kubernetes

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill model-serving-kubernetes -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills model-serving-kubernetes --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/model-serving-kubernetes .claude/skills/model-serving-kubernetes && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-serving-kubernetes
GitHub stars
47k
Used in
1 other repo
Token cost
~2.3k tokens
SKILL.md length
322 words
Files
1
Skills in repo
1,354
Repo updated
First seen
Licence
MIT

At a glance

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

  • Tasks that involve Container orchestration
  • SKILL.md covers When to Use This Skill, Prerequisites, KServe Installation and Basic InferenceService (KServe), plus 10 more sections
  • Calls kubectl, curl and helm; reaches kserve.github.io and prometheus-server.monitoring; needs HUGGING_FACE_HUB_TOKEN
  • Tasks that involve LLM inference and serving

What it does

Model Serving Kubernetes is an agent skill from sickn33/agentic-awesome-skills. Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and…

It sits in DevOps & Cloud, covering Container orchestration, LLM inference and serving and Machine learning. It works with Kubernetes and NVIDIA AI Platform. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Container orchestration
  • Tasks that involve LLM inference and serving
  • Tasks that involve Machine learning

Example prompts

  • “/model-serving-kubernetes”

Requirements

  • A credential in HUGGING_FACE_HUB_TOKEN
  • Compatibility (from SKILL.md): Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.

What it can do on your machine

Read from SKILL.md and the folder at commit ec02547. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • curl
    • helm
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • kserve.github.io
    • prometheus-server.monitoring

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HUGGING_FACE_HUB_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.

    From compatibility in the SKILL.md frontmatter.

Context cost

Model Serving Kubernetes loads about 2.3k tokens when it runs. Until then it costs about 31 tokens; SKILL.md has 322 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~31
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit ec02547, republished under its MIT licence (© sickn33). 322 words, ~2,318 tokens.

Download SKILL.mdSave it as .claude/skills/model-serving-kubernetes/SKILL.md (or your agent's skills folder).
name
model-serving-kubernetes
description
Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.
compatibility
Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.
category
devops
risk
critical
source
https://github.com/BagelHole/DevOps-Security-Agent-Skills
source_repo
BagelHole/DevOps-Security-Agent-Skills
source_type
community
date_added
2026-09-20
license
MIT
license_source
https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
metadata.author
devops-skills
metadata.version
1.0

Model Serving on Kubernetes

Production ML model serving with KServe and Triton — canary deployments, autoscaling, and GPU-aware scheduling.

When to Use This Skill

Use this skill when:

  • Serving scikit-learn, PyTorch, TensorFlow, or ONNX models at scale
  • Implementing canary deployments and A/B testing for ML models
  • Autoscaling inference pods based on request rate or GPU metrics
  • Deploying LLMs with Triton or KServe on Kubernetes
  • Managing multiple model versions with traffic splitting

Prerequisites

  • Kubernetes 1.28+ with GPU nodes
  • KServe installed (or Triton standalone)
  • kubectl and helm configured
  • NVIDIA GPU Operator installed on cluster

KServe Installation

bash
# Install KServe with Helm
helm repo add kserve https://kserve.github.io/helm-charts
helm repo update

helm install kserve kserve/kserve \
  --namespace kserve \
  --create-namespace \
  --set kserve.controller.gateway.ingressGateway.className=nginx

# Verify
kubectl get pods -n kserve
kubectl get crd | grep kserve

Basic InferenceService (KServe)

yaml
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
  name: sklearn-iris
  namespace: models
spec:
  predictor:
    sklearn:
      storageUri: gs://kfserving-examples/models/sklearn/1.0/model
      resources:
        requests:
          cpu: "1"
          memory: 2Gi
        limits:
          cpu: "2"
          memory: 4Gi
bash
kubectl apply -f inference-service.yaml

# Get inference service URL
kubectl get inferenceservice sklearn-iris -n models
# NAME           URL                                          READY   ...
# sklearn-iris   http://sklearn-iris.models.example.com       True

# Test prediction
curl -X POST http://sklearn-iris.models.example.com/v1/models/sklearn-iris:predict \
  -H "Content-Type: application/json" \
  -d '{"instances": [[6.8, 2.8, 4.8, 1.4]]}'

GPU-Enabled LLM InferenceService

yaml
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
  name: llama-3-8b
  namespace: models
  annotations:
    serving.kserve.io/enable-prometheus-scraping: "true"
spec:
  predictor:
    containers:
    - name: vllm-container
      image: vllm/vllm-openai:latest
      args:
      - "--model"
      - "meta-llama/Llama-3.1-8B-Instruct"
      - "--tensor-parallel-size"
      - "1"
      - "--gpu-memory-utilization"
      - "0.90"
      ports:
      - containerPort: 8080
        protocol: TCP
      resources:
        requests:
          nvidia.com/gpu: "1"
          memory: "20Gi"
          cpu: "4"
        limits:
          nvidia.com/gpu: "1"
          memory: "24Gi"
          cpu: "8"
      readinessProbe:
        httpGet:
          path: /health
          port: 8080
        initialDelaySeconds: 60
        periodSeconds: 10
      env:
      - name: HUGGING_FACE_HUB_TOKEN
        valueFrom:
          secretKeyRef:
            name: hf-token
            key: token
    nodeSelector:
      nvidia.com/gpu.present: "true"
  transformer:
    containers:
    - name: kserve-container
      image: kserve/kserve-transformer:latest

Canary Deployment (Traffic Splitting)

yaml
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
  name: llama-3-8b
  namespace: models
spec:
  predictor:
    canaryTrafficPercent: 20    # 20% to new version, 80% to stable
    containers:
    - name: vllm-container
      image: vllm/vllm-openai:latest
      args:
      - "--model"
      - "meta-llama/Llama-3.1-8B-Instruct-v2"  # new model version
      resources:
        limits:
          nvidia.com/gpu: "1"
bash
# Gradually increase canary traffic
kubectl patch inferenceservice llama-3-8b -n models \
  --type='json' \
  -p='[{"op":"replace","path":"/spec/predictor/canaryTrafficPercent","value":50}]'

# Promote canary to stable
kubectl patch inferenceservice llama-3-8b -n models \
  --type='json' \
  -p='[{"op":"remove","path":"/spec/predictor/canaryTrafficPercent"}]'

Autoscaling with KEDA

yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: llama-scaler
  namespace: models
spec:
  scaleTargetRef:
    apiVersion: serving.kserve.io/v1beta1
    kind: InferenceService
    name: llama-3-8b
  minReplicaCount: 1
  maxReplicaCount: 5
  triggers:
  - type: prometheus
    metadata:
      serverAddress: http://prometheus-server.monitoring:9090
      metricName: kserve_request_count
      threshold: "10"
      query: |
        sum(rate(kserve_request_count_total{namespace="models",
            service="llama-3-8b"}[1m]))

NVIDIA Triton Inference Server

yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: triton-server
  namespace: models
spec:
  replicas: 2
  selector:
    matchLabels:
      app: triton
  template:
    metadata:
      labels:
        app: triton
    spec:
      containers:
      - name: triton
        image: nvcr.io/nvidia/tritonserver:24.05-py3
        args:
        - "tritonserver"
        - "--model-store=s3://my-model-store/models"
        - "--model-control-mode=poll"        # auto-load new model versions
        - "--repository-poll-secs=30"
        - "--metrics-port=8002"
        ports:
        - containerPort: 8000   # HTTP
        - containerPort: 8001   # gRPC
        - containerPort: 8002   # Metrics
        resources:
          limits:
            nvidia.com/gpu: "1"
        readinessProbe:
          httpGet:
            path: /v2/health/ready
            port: 8000
          initialDelaySeconds: 30

Triton Model Repository Structure

s3://my-model-store/models/
├── text-classifier/
│   ├── config.pbtxt
│   ├── 1/
│   │   └── model.onnx
│   └── 2/
│       └── model.onnx          # new version; auto-loaded
├── embedding-model/
│   ├── config.pbtxt
│   └── 1/
│       └── model.onnx
protobuf
# config.pbtxt for ONNX model
name: "text-classifier"
backend: "onnxruntime"
max_batch_size: 64
dynamic_batching {
  preferred_batch_size: [16, 32]
  max_queue_delay_microseconds: 1000
}
input [
  { name: "input_ids" data_type: TYPE_INT64 dims: [-1] }
  { name: "attention_mask" data_type: TYPE_INT64 dims: [-1] }
]
output [
  { name: "logits" data_type: TYPE_FP32 dims: [-1] }
]
instance_group [
  { kind: KIND_GPU count: 2 }   # 2 model instances on GPU
]

Model Management Commands

bash
# List loaded models (Triton)
curl http://triton:8000/v2/models

# Load a new model version
curl -X POST http://triton:8000/v2/repository/models/text-classifier/load

# Unload a model
curl -X POST http://triton:8000/v2/repository/models/text-classifier/unload

# KServe — watch rollout status
kubectl rollout status deployment/llama-3-8b-predictor -n models
kubectl get inferenceservice llama-3-8b -n models -w

Common Issues

IssueCauseFix
InferenceService not readyModel loading or OOMCheck predictor pod logs; increase memory limits
Canary stuck at 0%KNative routing issueCheck kubectl get ksvc -n models
Triton missing modelS3 permissions or pathVerify IAM role; check --model-store path
Low GPU utilizationDynamic batching offEnable dynamic_batching in Triton config
Autoscaler not triggeringPrometheus query wrongTest query in Prometheus UI

Best Practices

  • Use canary deployments for all model updates — roll back in seconds if metrics degrade.
  • Enable Triton dynamic batching — it can increase GPU throughput 5–10× for small models.
  • Store models in S3/GCS with versioned paths (s3://bucket/model/v1/, v2/).
  • Pin GPU node selectors to prevent model pods landing on CPU-only nodes.
  • Monitor p99 latency and error rates per model version during canary rollouts.
  • vllm-server (vllm-server) - vLLM for LLM serving
  • llm-inference-scaling (llm-inference-scaling) - KEDA autoscaling
  • kubernetes-ops (kubernetes-ops) - Core Kubernetes operations
  • gpu-server-management (gpu-server-management) - GPU nodes

Limitations

  • Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
  • Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.
Example
bash
git status && git diff --stat
kubectl diff -f manifest.yaml

Adapted from BagelHole/DevOps-Security-Agent-Skills (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/model-serving-kubernetes of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit ec02547

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Model Serving Kubernetes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Serving Kubernetes compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Serving Kubernetes this skillsickn33/agentic-awesome-skills47k1 repos~2.3kAutomated safety check: PassMIT
Model Serving Kubernetesmajiayu000/claude-skill-registry6661 repos~2.1kAutomated safety check: PassMIT
Dstack Presetsdstackai/dstack2.3k—~403Automated safety check: PassMPL-2.0
Deploy Controllerai-runway/airunway101—~927Automated safety check: PassApache-2.0
Gke Manifest Generationgoogle/skills21k—~3.1kAutomated safety check: PassApache-2.0
Dynamo Recipe RunnerNVIDIA/skills3.5k—~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • Model Serving Kubernetes

    majiayu000/claude-skill-registry

    Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

    666 GitHub starsUsed in 1 repo~2.1k tokens
    DevOps & CloudAuto-check passed
  • Dstack Presets

    dstackai/dstack

    Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.

    2.3k GitHub stars~403 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Deploy Controller

    ai-runway/airunway

    Interactively build, push or load, and deploy an airunway component (controller or any provider) to the cluster

    101 GitHub stars~927 tokensUpdated 11 days ago
    DevOps & CloudAuto-check passed
  • Official

    Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters.

    21k GitHub stars~3.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Official

    Select, validate, patch, and deploy existing NVIDIA Dynamo Kubernetes recipes.

    3.5k GitHub stars~1.8k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Official

    A skill your agent uses when the user is hands-on deploying an in-bundle DOCA service container (Argus, DMS, Firefly, or UROM service) on a BlueField — kubelet standalone watching a static-pod…

    3.5k GitHub stars~2.5k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,354 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Content Creator

    sickn33/agentic-awesome-skills

    Drafts and reviews audience-specific content from supplied brand examples, with local scripts for brand voice and SEO diagnostics, channel templates and a content calendar.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed

Questions about Model Serving Kubernetes

What does Model Serving Kubernetes do?

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. Model Serving Kubernetes is an agent skill from sickn33/agentic-awesome-skills. Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

When should I use Model Serving Kubernetes?

Model Serving Kubernetes fits situations like: tasks that involve Container orchestration; tasks that involve LLM inference and serving; tasks that involve Machine learning.

How do I install Model Serving Kubernetes in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill model-serving-kubernetes -a claude-code`. Or copy the skill folder (skills/model-serving-kubernetes in sickn33/agentic-awesome-skills) into .claude/skills/model-serving-kubernetes in your project. Claude Code loads it when a task matches its description.

How do I install Model Serving Kubernetes in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill model-serving-kubernetes -a codex`. Or copy the skill folder (skills/model-serving-kubernetes in sickn33/agentic-awesome-skills) into .agents/skills/model-serving-kubernetes in your project. Codex loads it when a task matches its description.

Can I use Model Serving Kubernetes in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill model-serving-kubernetes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-serving-kubernetes, .gemini/skills/model-serving-kubernetes, .github/skills/model-serving-kubernetes and .opencode/skills/model-serving-kubernetes in your project.

What does Model Serving Kubernetes need to run?

Going by SKILL.md and its folder, Model Serving Kubernetes needs the command-line tools its instructions call (kubectl, curl, helm and git) and credentials named HUGGING_FACE_HUB_TOKEN. Our summary lists: A credential in HUGGING_FACE_HUB_TOKEN. Compatibility (from SKILL.md): Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled..

Does Model Serving Kubernetes access the network?

SKILL.md names 3 domains. In commands or code: kserve.github.io and prometheus-server.monitoring; the agent is likely to contact these when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Model Serving Kubernetes safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Model Serving Kubernetes use?

Model Serving Kubernetes is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Serving Kubernetes use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Model Serving Kubernetes?

Skills that share tags, products or a category with Model Serving Kubernetes: Model Serving Kubernetes (majiayu000/claude-skill-registry, 666 stars), Dstack Presets (dstackai/dstack, 2.3k stars), Deploy Controller (ai-runway/airunway, 101 stars) and Gke Manifest Generation (google/skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Serving Kubernetes?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,343 GitHub stars. The repository holds 1,354 skills in this directory. The repository was last updated on October 7, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.