Official agent skill

Gke Inference

by google in google/skills

Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Gke Inference

skills CLI
$ npx skills add google/skills --skill gke-inference -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills gke-inference --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-inference .claude/skills/gke-inference && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gke-inference
GitHub stars
21k
Token cost
~2k tokens
SKILL.md length
492 words
Files
1
Skills in repo
145
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.

  • Works in 3 steps: Discovery: Find Models and Hardware → Generate Manifest → Review and Deploy
  • Deploying GKE inference servers
  • SKILL.md covers When to Use, Prerequisites, Workflow and GPU ComputeClass for Inference, plus 4 more sections
  • Calls gcloud and kubectl

What it does

Gke Inference is an agent skill from google/skills, published by the product's own GitHub organization. Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. Use when deploying GKE inference servers, configuring GKE GPU resources for inference, or deploying LLMs on GKE. Don't use for migrating existing AI workloads to GKE (use google-cloud-solution-guided-gke-ai-migration), GKE RAG with Cloud SQL/AlloyDB (use google-cloud-solution-rag-enterprise-search-gke-sqldb), or batch/HPC (use gke-batch-hpc).

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Retrieval-augmented generation, LLM inference and serving and SQL. It works with Google Kubernetes Engine, Google Cloud, SQL and Kubernetes. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • Deploying GKE inference servers
  • Configuring GKE GPU resources for inference
  • Deploying LLMs on GKE
  • Migrating existing AI workloads to GKE (use google-cloud-solution-guided-gke-ai-migration)

Example prompts

  • “Use the gke-inference skill to deploy and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers”
  • “/gke-inference”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Discovery: Find Models and Hardware
  2. Generate Manifest
  3. Review and Deploy

What it can do on your machine

Read from SKILL.md and the folder at commit 8a1ac05. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gcloud
    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gcloud and kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gke Inference loads about 2k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 492 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from google/skills at commit 8a1ac05, republished under its Apache-2.0 licence (© google). 492 words, ~1,988 tokens.

Download SKILL.mdSave it as .claude/skills/gke-inference/SKILL.md (or your agent's skills folder).
name
gke-inference
description
Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. Use when deploying GKE inference servers, configuring GKE GPU resources for inference, or deploying LLMs on GKE. Don't use for migrating existing AI workloads to GKE (use google-cloud-solution-guided-gke-ai-migration), GKE RAG with Cloud SQL/AlloyDB (use google-cloud-solution-rag-enterprise-search-gke-sqldb), or batch/HPC (use gke-batch-hpc).
metadata.version
1.0.1
metadata.category
Containers

GKE AI/ML Inference

Routing Note: For migrating existing AI workloads to GKE, open google-cloud-solution-guided-gke-ai-migration/SKILL.md. For GKE RAG with Cloud SQL or AlloyDB (pgvector), open google-cloud-solution-rag-enterprise-search-gke-sqldb/SKILL.md.

This reference covers deploying AI/ML inference workloads on GKE using Google's Inference Quickstart (GIQ) and best practices for LLM serving.

MCP Tools: apply_k8s_manifest, get_k8s_resource, get_k8s_logs, get_k8s_rollout_status, describe_k8s_resource, list_k8s_events. CLI-only: gcloud container ai profiles *

When to Use

  • Deploy an AI model (Llama, Gemma, Mistral, etc.) to GKE
  • Generate optimized Kubernetes manifests for inference
  • Select GPU/TPU accelerators for model serving
  • Configure autoscaling for LLM inference

Prerequisites

  • A golden path GKE Autopilot cluster (GPU workloads are supported via ComputeClasses and NAP)
  • gcloud CLI authenticated
  • Sufficient GPU/TPU quota in the target region

Workflow

1. Discovery: Find Models and Hardware
bash
# List all supported models
gcloud container ai profiles models list --quiet

# Find valid accelerator/server combinations for a model
gcloud container ai profiles list --model=<MODEL_NAME> --quiet

# Example: what can run Gemma 2 9B?
gcloud container ai profiles list --model=gemma-2-9b-it --quiet
2. Generate Manifest
bash
gcloud container ai profiles manifests create \
  --model=<MODEL_NAME> \
  --model-server=<SERVER> \
  --accelerator-type=<ACCELERATOR> \
  --target-ntpot-milliseconds=<NTPOT> --quiet > inference.yaml

Parameters:

  • --model: Model ID (e.g., gemma-2-9b-it, llama-3-8b)
  • --model-server: Inference server (vllm, tgi, triton, tensorrt-llm)
  • --accelerator-type: GPU/TPU type (nvidia-l4, nvidia-tesla-a100, nvidia-h100-80gb)
  • --target-ntpot-milliseconds: Target Normalized Time Per Output Token (optional, for latency optimization)

Example:

bash
gcloud container ai profiles manifests create \
  --model=gemma-2-9b-it \
  --model-server=vllm \
  --accelerator-type=nvidia-l4 \
  --target-ntpot-milliseconds=50 --quiet > inference.yaml
3. Review and Deploy
bash
# Review for placeholders (HF tokens, PVCs)
cat inference.yaml

# Deploy
kubectl apply -f inference.yaml

# Monitor
kubectl get pods -w
kubectl logs -f <POD_NAME>

Some models require Hugging Face tokens. Create a Kubernetes Secret and reference it in the manifest.

GPU ComputeClass for Inference

For Autopilot clusters, create a ComputeClass to target GPU nodes:

yaml
apiVersion: cloud.google.com/v1
kind: ComputeClass
metadata:
  name: l4-inference
spec:
  priorities:
  - machineFamily: g2
    gpu:
      type: nvidia-l4
      count: 1
    minCores: 4
    minMemoryGb: 16

Accelerator Selection Guide

AcceleratorBest ForMemoryRelative Cost
NVIDIA T4Budget inference,16 GBLowest
: : lightweight legacy : : :
: : models : : :
NVIDIA L4 (G2)Small-medium model24 GBLow
: : inference, video, : : :
: : graphics : : :
NVIDIA RTX PRO 6000Multimodal AI,96 GBMedium
: (G4) : high-fidelity 3D, : : :
: : fine-tuning : : :
Cloud TPU v5eCost-effectiveVariesMedium
: : transformer inference : : :
Cloud TPU v5pHigh-performanceVariesHigh
: : training : : :
Cloud TPU v6eHigh-efficiency next-gen32 GB/chipMedium-High
: (Trillium) : training & serving : : :
Cloud TPU v7xUltra-scale inference &192 GB/chipHigh
: (Ironwood) : agentic workflows : : :
NVIDIA A100Large model inference,40/80 GBHigh
: : enterprise ML : : :
NVIDIA H100 / H200Frontier model training,80/141 GBHighest
: : high throughput : : :
NVIDIA B200 (A4)Blackwell-scale192 GBHighest
: : training, FP4 precision : : :
NVIDIA GB200 (A4X)Rack-scale AI (GraceMassiveHighest
: : Blackwell Superchip) : : :
Show full SKILL.md (182 more words)Show less

Autoscaling LLM Inference

GPU-based autoscaling

Use custom metrics for GPU utilization:

yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: llm-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: llm-server
  minReplicas: 1
  maxReplicas: 10
  metrics:
  - type: Pods
    pods:
      metric:
        name: gpu_duty_cycle
      target:
        type: AverageValue
        averageValue: "80"
Best practices for inference autoscaling
  1. Use DCGM metrics: Golden path enables DCGM monitoring for GPU utilization metrics
  2. Set appropriate minReplicas: At least 1 for always-on serving; 0 for batch/on-demand
  3. Tune scale-down delay: LLM model loading is slow; use longer stabilization windows
  4. Consider queue depth: Scale on pending requests rather than pure GPU utilization for latency-sensitive workloads

Optimization Tips

  • Quantization: Use quantized models (GPTQ, AWQ) to reduce GPU memory and increase throughput
  • Batching: Configure model server batch size for throughput vs latency trade-off
  • Tensor parallelism: Split large models across multiple GPUs within a node
  • KV cache optimization: Tune --gpu-memory-utilization in vLLM for KV cache allocation

Troubleshooting

IssueCauseFix
InvalidUnsupported tupleRe-run `gcloud container ai
: model/accelerator : : profiles list :
: combination : : --model=<MODEL>` :
GPU quota exceededRegional quota limitRequest quota increase or
: : : try a different region :
OOM on GPUModel too large forUse larger GPU, enable
: : accelerator : quantization, or use tensor :
: : : parallelism :
Slow cold startLarge model loading fromUse local SSD for model
: : registry : caching; pre-pull images :

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/cloud/gke-inference of google/skills.

Open the folder on GitHubat commit 8a1ac05

Compare with similar skills

Gke Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gke Inference compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gke Inference this skillgoogle/skills21k—~2kAutomated safety check: PassApache-2.0
Dd GCP Integrationdatadog-labs/agent-skills177—~8kAutomated safety check: NotesMIT
GCP ArchitectFerroxLabs/wayland608—~4.4kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Query Finelogmarin-community/marin3.9k—~1.1kAutomated safety check: PassApache-2.0
Searching DocumentsGAIK-project/gaik-toolkit100—~4.2kAutomated safety check: PassMIT

Similar skills

  • Dd GCP Integration

    datadog-labs/agent-skills

    Set up the Datadog Google Cloud integration with Terraform - creates a service account in the host project, lets Datadog's delegate principal impersonate it via roles/iam.serviceAccountTokenCreator…

    177 GitHub stars~8k tokensUpdated 5 days ago
    DevOps & CloudAuto-check: notes
  • GCP Architect

    FerroxLabs/wayland

    GCP architecture. An agent skill from FerroxLabs/wayland.

    608 GitHub stars~4.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Query Finelog

    marin-community/marin

    Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding.

    3.9k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Searching Documents

    GAIK-project/gaik-toolkit

    Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…

    100 GitHub stars~4.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Aliyun Opensearch Search

    cinience/alicloud-skills

    A skill your agent uses when working with OpenSearch vector search edition via the Python SDK (ha3engine) to push documents and run HA/SQL searches.

    397 GitHub stars~1.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from google/skills

All 145 skills in this repo
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated today
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Official

    Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans.

    21k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Questions about Gke Inference

What does Gke Inference do?

Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. Gke Inference is an agent skill from google/skills, published by the product's own GitHub organization. Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.

When should I use Gke Inference?

Gke Inference fits situations like: deploying GKE inference servers; configuring GKE GPU resources for inference; deploying LLMs on GKE; migrating existing AI workloads to GKE (use google-cloud-solution-guided-gke-ai-migration).

How do I install Gke Inference in Claude Code?

Run `npx skills add google/skills --skill gke-inference -a claude-code`. Or copy the skill folder (skills/cloud/gke-inference in google/skills) into .claude/skills/gke-inference in your project. Claude Code loads it when a task matches its description.

How do I install Gke Inference in Codex?

Run `npx skills add google/skills --skill gke-inference -a codex`. Or copy the skill folder (skills/cloud/gke-inference in google/skills) into .agents/skills/gke-inference in your project. Codex loads it when a task matches its description.

Can I use Gke Inference in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-inference, .gemini/skills/gke-inference, .github/skills/gke-inference and .opencode/skills/gke-inference in your project.

What does Gke Inference need to run?

Going by SKILL.md and its folder, Gke Inference needs the command-line tools its instructions call (gcloud and kubectl).

Does Gke Inference access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gke Inference safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gke Inference use?

Gke Inference is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gke Inference use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gke Inference?

Skills that share tags, products or a category with Gke Inference: Dd GCP Integration (datadog-labs/agent-skills, 177 stars), GCP Architect (FerroxLabs/wayland, 608 stars), Dstack Prototyping (dstackai/dstack, 2.3k stars) and Query Finelog (marin-community/marin, 3.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gke Inference?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 20,994 GitHub stars. The repository holds 145 skills in this directory. The repository was last updated on October 6, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.