Dd GCP Integration
datadog-labs/agent-skills
Set up the Datadog Google Cloud integration with Terraform - creates a service account in the host project, lets Datadog's delegate principal impersonate it via roles/iam.serviceAccountTokenCreator…
Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.
$ npx skills add google/skills --skill gke-inference -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills gke-inference --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-inference .claude/skills/gke-inference && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gke-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-inference into .claude/skills/gke-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-inference", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/gke-inferenceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill gke-inference -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills gke-inference --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/gke-inference .agents/skills/gke-inference && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gke-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-inference into .agents/skills/gke-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-inference", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-inference -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills gke-inference --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/gke-inference .cursor/skills/gke-inference && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gke-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-inference into .cursor/skills/gke-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-inference", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/gke-inference--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill gke-inference -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills gke-inference --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/gke-inference .gemini/skills/gke-inference && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gke-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-inference into .gemini/skills/gke-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-inference", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills gke-inferenceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill gke-inference -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/gke-inference .github/skills/gke-inference && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gke-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-inference into .github/skills/gke-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-inference", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-inference -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills gke-inference --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/gke-inference .opencode/skills/gke-inference && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gke-inference" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-inference into .opencode/skills/gke-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-inference", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gke-inferenceDeploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.
Gke Inference is an agent skill from google/skills, published by the product's own GitHub organization. Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. Use when deploying GKE inference servers, configuring GKE GPU resources for inference, or deploying LLMs on GKE. Don't use for migrating existing AI workloads to GKE (use google-cloud-solution-guided-gke-ai-migration), GKE RAG with Cloud SQL/AlloyDB (use google-cloud-solution-rag-enterprise-search-gke-sqldb), or batch/HPC (use gke-batch-hpc).
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Retrieval-augmented generation, LLM inference and serving and SQL. It works with Google Kubernetes Engine, Google Cloud, SQL and Kubernetes. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8a1ac05. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gcloudkubectlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gcloud and kubectl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gke Inference loads about 2k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 492 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from google/skills at commit 8a1ac05, republished under its Apache-2.0 licence (© google). 492 words, ~1,988 tokens.
.claude/skills/gke-inference/SKILL.md (or your agent's skills folder).Routing Note: For migrating existing AI workloads to GKE, open
google-cloud-solution-guided-gke-ai-migration/SKILL.md. For GKE RAG with Cloud SQL or AlloyDB (pgvector), opengoogle-cloud-solution-rag-enterprise-search-gke-sqldb/SKILL.md.
This reference covers deploying AI/ML inference workloads on GKE using Google's Inference Quickstart (GIQ) and best practices for LLM serving.
MCP Tools:
apply_k8s_manifest,get_k8s_resource,get_k8s_logs,get_k8s_rollout_status,describe_k8s_resource,list_k8s_events. CLI-only:gcloud container ai profiles *
gcloud CLI authenticated# List all supported models
gcloud container ai profiles models list --quiet
# Find valid accelerator/server combinations for a model
gcloud container ai profiles list --model=<MODEL_NAME> --quiet
# Example: what can run Gemma 2 9B?
gcloud container ai profiles list --model=gemma-2-9b-it --quietgcloud container ai profiles manifests create \
--model=<MODEL_NAME> \
--model-server=<SERVER> \
--accelerator-type=<ACCELERATOR> \
--target-ntpot-milliseconds=<NTPOT> --quiet > inference.yamlParameters:
--model: Model ID (e.g., gemma-2-9b-it, llama-3-8b)--model-server: Inference server (vllm, tgi, triton, tensorrt-llm)--accelerator-type: GPU/TPU type (nvidia-l4, nvidia-tesla-a100,
nvidia-h100-80gb)--target-ntpot-milliseconds: Target Normalized Time Per Output Token
(optional, for latency optimization)Example:
gcloud container ai profiles manifests create \
--model=gemma-2-9b-it \
--model-server=vllm \
--accelerator-type=nvidia-l4 \
--target-ntpot-milliseconds=50 --quiet > inference.yaml# Review for placeholders (HF tokens, PVCs)
cat inference.yaml
# Deploy
kubectl apply -f inference.yaml
# Monitor
kubectl get pods -w
kubectl logs -f <POD_NAME>Some models require Hugging Face tokens. Create a Kubernetes Secret and reference it in the manifest.
For Autopilot clusters, create a ComputeClass to target GPU nodes:
apiVersion: cloud.google.com/v1
kind: ComputeClass
metadata:
name: l4-inference
spec:
priorities:
- machineFamily: g2
gpu:
type: nvidia-l4
count: 1
minCores: 4
minMemoryGb: 16| Accelerator | Best For | Memory | Relative Cost |
|---|---|---|---|
| NVIDIA T4 | Budget inference, | 16 GB | Lowest |
| : : lightweight legacy : : : | |||
| : : models : : : | |||
| NVIDIA L4 (G2) | Small-medium model | 24 GB | Low |
| : : inference, video, : : : | |||
| : : graphics : : : | |||
| NVIDIA RTX PRO 6000 | Multimodal AI, | 96 GB | Medium |
| : (G4) : high-fidelity 3D, : : : | |||
| : : fine-tuning : : : | |||
| Cloud TPU v5e | Cost-effective | Varies | Medium |
| : : transformer inference : : : | |||
| Cloud TPU v5p | High-performance | Varies | High |
| : : training : : : | |||
| Cloud TPU v6e | High-efficiency next-gen | 32 GB/chip | Medium-High |
| : (Trillium) : training & serving : : : | |||
| Cloud TPU v7x | Ultra-scale inference & | 192 GB/chip | High |
| : (Ironwood) : agentic workflows : : : | |||
| NVIDIA A100 | Large model inference, | 40/80 GB | High |
| : : enterprise ML : : : | |||
| NVIDIA H100 / H200 | Frontier model training, | 80/141 GB | Highest |
| : : high throughput : : : | |||
| NVIDIA B200 (A4) | Blackwell-scale | 192 GB | Highest |
| : : training, FP4 precision : : : | |||
| NVIDIA GB200 (A4X) | Rack-scale AI (Grace | Massive | Highest |
| : : Blackwell Superchip) : : : |
Use custom metrics for GPU utilization:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: llm-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: llm-server
minReplicas: 1
maxReplicas: 10
metrics:
- type: Pods
pods:
metric:
name: gpu_duty_cycle
target:
type: AverageValue
averageValue: "80"--gpu-memory-utilization in vLLM for KV
cache allocation| Issue | Cause | Fix |
|---|---|---|
| Invalid | Unsupported tuple | Re-run `gcloud container ai |
| : model/accelerator : : profiles list : | ||
: combination : : --model=<MODEL>` : | ||
| GPU quota exceeded | Regional quota limit | Request quota increase or |
| : : : try a different region : | ||
| OOM on GPU | Model too large for | Use larger GPU, enable |
| : : accelerator : quantization, or use tensor : | ||
| : : : parallelism : | ||
| Slow cold start | Large model loading from | Use local SSD for model |
| : : registry : caching; pre-pull images : |
© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/cloud/gke-inference of google/skills.
Open the folder on GitHubat commit 8a1ac05
Gke Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gke Inference this skillgoogle/skills | 21k | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Dd GCP Integrationdatadog-labs/agent-skills | 177 | — | ~8k | Automated safety check: Notes | MIT | |
| GCP ArchitectFerroxLabs/wayland | 608 | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| Query Finelogmarin-community/marin | 3.9k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Searching DocumentsGAIK-project/gaik-toolkit | 100 | — | ~4.2k | Automated safety check: Pass | MIT |
datadog-labs/agent-skills
Set up the Datadog Google Cloud integration with Terraform - creates a service account in the host project, lets Datadog's delegate principal impersonate it via roles/iam.serviceAccountTokenCreator…
FerroxLabs/wayland
GCP architecture. An agent skill from FerroxLabs/wayland.
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
marin-community/marin
Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding.
GAIK-project/gaik-toolkit
Builds and debugs retrieval with the gaik toolkit — PgVectorStore, Ranker, FinnishTextProcessor, RelevanceGate — as hybrid search: pgvector similarity plus Postgres full-text, fused by rank, and the…
cinience/alicloud-skills
A skill your agent uses when working with OpenSearch vector search edition via the Python SDK (ha3engine) to push documents and run HA/SQL searches.
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
google/skills
Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans.
Categories
Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. Gke Inference is an agent skill from google/skills, published by the product's own GitHub organization. Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.
Gke Inference fits situations like: deploying GKE inference servers; configuring GKE GPU resources for inference; deploying LLMs on GKE; migrating existing AI workloads to GKE (use google-cloud-solution-guided-gke-ai-migration).
Run `npx skills add google/skills --skill gke-inference -a claude-code`. Or copy the skill folder (skills/cloud/gke-inference in google/skills) into .claude/skills/gke-inference in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill gke-inference -a codex`. Or copy the skill folder (skills/cloud/gke-inference in google/skills) into .agents/skills/gke-inference in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-inference, .gemini/skills/gke-inference, .github/skills/gke-inference and .opencode/skills/gke-inference in your project.
Going by SKILL.md and its folder, Gke Inference needs the command-line tools its instructions call (gcloud and kubectl).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gke Inference is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Gke Inference: Dd GCP Integration (datadog-labs/agent-skills, 177 stars), GCP Architect (FerroxLabs/wayland, 608 stars), Dstack Prototyping (dstackai/dstack, 2.3k stars) and Query Finelog (marin-community/marin, 3.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 20,994 GitHub stars. The repository holds 145 skills in this directory. The repository was last updated on October 6, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.