Agent skill

Gke Inference Stack Deploy

by GoogleCloudPlatform in GoogleCloudPlatform/accelerated-platforms

Deploys the GKE base platform and inference-specific terra-services (GPU/TPU) for accelerated workloads.

Apache-2.0Auto-check passedDevOps & Cloud

Install Gke Inference Stack Deploy

skills CLI
$ npx skills add GoogleCloudPlatform/accelerated-platforms --skill gke-inference-stack-deploy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GoogleCloudPlatform/accelerated-platforms gke-inference-stack-deploy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GoogleCloudPlatform/accelerated-platforms.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gke-inference-stack-deploy .claude/skills/gke-inference-stack-deploy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gke-inference-stack-deploy
GitHub stars
106
Token cost
~1.1k tokens
SKILL.md length
296 words
Files
2
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deploys the GKE base platform and inference-specific terra-services (GPU/TPU) for accelerated workloads.

  • Works in 5 steps: Prerequisites & Platform Configuration → Accelerator & Hardware Configuration → Deploy Inference Platform & Prerequisite… → …
  • DevOps & Cloud work in your project
  • SKILL.md covers 1. Prerequisites & Platform…, 2. Accelerator & Hardware…, 3. Deploy Inference Platform &… and 4. Deploy Inference…, plus 1 more section
  • Calls terraform, kubectl and gcloud

What it does

Gke Inference Stack Deploy is an agent skill from GoogleCloudPlatform/accelerated-platforms. Deploys the GKE base platform and inference-specific terra-services (GPU/TPU) for accelerated workloads.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/evals.json`).

It sits in DevOps & Cloud. It works with Google Kubernetes Engine and Google Cloud. The repository describes itself as: This repository is a collection of accelerated platform best practices, reference architectures, example use cases, reference implementations, and various other assets on Google… The licence is Apache-2.0.

When your agent uses it

  • DevOps & Cloud work in your project

Example prompts

  • “Use the gke-inference-stack-deploy skill to deploy the GKE base platform and inference-specific terra-services (GPU/TPU) for accelerated workloads”
  • “/gke-inference-stack-deploy”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): kubectl, gcloud, terraform, python3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Prerequisites & Platform Configuration
  2. Accelerator & Hardware Configuration
  3. Deploy Inference Platform & Prerequisite Terra-Services
  4. Deploy Inference Terra-Services
  5. Verification

What it can do on your machine

Read from SKILL.md and the folder at commit bab190a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • kubectl
    • gcloud
    • terraform
    • python3

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • terraform
    • kubectl
    • gcloud

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl and gcloud, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gke Inference Stack Deploy loads about 1.1k tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 296 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GoogleCloudPlatform/accelerated-platforms at commit bab190a, republished under its Apache-2.0 licence (© GoogleCloudPlatform). 296 words, ~1,066 tokens.

Download SKILL.mdSave it as .claude/skills/gke-inference-stack-deploy/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
gke-inference-stack-deploy
description
Deploys the GKE base platform and inference-specific terra-services (GPU/TPU) for accelerated workloads.
allowed-tools
kubectl, gcloud, terraform, python3
version
0.1.0

Deploy GKE Inference Stack Skill

Follow these instructions to provision the GKE Base Platform and inference-specific terra-services.

1. Prerequisites & Platform Configuration

  1. Ask the user (only for information not provided in the user's prompt):

    • "What is the path to the repository?" -> Set ACP_REPO_DIR
    • "What is your Google Cloud project ID?" -> Set platform_default_project_id
    • "What GKE region would you like to use?" -> Set cluster_region
    • "What platform name would you like to set?" -> Set platform_name
    • "Would you like to deploy an Autopilot or Standard cluster?" -> Set cluster_type
  2. Action: Inject these values into the appropriate tfvars files (${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars and cluster.auto.tfvars) using sed.

    bash
    # Update platform variables
    sed -i 's/^platform_name.*/platform_name = "<platform_name>"/g' "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars"
    grep -q "^platform_default_project_id" "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars" || echo "platform_default_project_id = \"\"" >> "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars"
    sed -i 's/^platform_default_project_id.*/platform_default_project_id = "<project_id>"/g' "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars"
    
    # Update cluster variables
    sed -i 's/^cluster_region.*/cluster_region = "<cluster_region>"/g' "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/cluster.auto.tfvars"

2. Accelerator & Hardware Configuration

  1. Ask the user: "Which accelerator(s) would you like to deploy for inference? You can select GPU, TPU, or both."

    • Options: GPU (rtx-pro-6000, h100, h200), TPU (v6e), or both.
  2. Action:

    • Determine if the target inference terra-services are online_gpu, online_tpu, or both based on the chosen accelerator(s).
    • Update accelerator configurations in the shared tfvars if necessary based on user selection.

3. Deploy Inference Platform & Prerequisite Terra-Services

Run the appropriate inference reference architecture deployment script based on the cluster type selected by the user. This provisions the core platform along with prerequisite inference services (such as Hugging Face and monitoring initialization). Always refer to existing scripts for execution rather than defining the CORE_TERRASERVICES_APPLY array explicitly.

  • For Standard Cluster:
    bash
    "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/terraform/deploy-standard.sh"
  • For Autopilot Cluster:
    bash
    "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/terraform/deploy-ap.sh"

4. Deploy Inference Terra-Services

Navigate to the specific inference terra-service directories (online_gpu, online_tpu, or both) and execute Terraform commands.

bash
# Depending on what the user chose, add "online_gpu" and/or "online_tpu" to this array
declare -a selected_terraservices=( "online_gpu" "online_tpu" )

for inference_terraservice in "${selected_terraservices[@]}"; do
  cd "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/terraform/${inference_terraservice}/"
  terraform init
  terraform plan -input=false -out=tfplan
  terraform apply -input=false tfplan
  rm tfplan
done

5. Verification

Retrieve GKE cluster credentials and verify the deployment. Note: if using NAP (Node Auto-Provisioning), nodes may not be visible until workloads are deployed.

bash
gcloud container clusters get-credentials "<platform_name>" --region "<cluster_region>" --project "<project_id>" --dns-endpoint
kubectl get computeclasses
kubectl get namespaces # You should see the online_gpu and/or online_tpu namespace depending on what was deployed

© GoogleCloudPlatform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/gke-inference-stack-deploy of GoogleCloudPlatform/accelerated-platforms.

  • SKILL.md
  • evals/evals.json

Open the folder on GitHubat commit bab190a

Compare with similar skills

Gke Inference Stack Deploy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gke Inference Stack Deploy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gke Inference Stack Deploy this skillGoogleCloudPlatform/accelerated-platforms106—~1.1kAutomated safety check: PassApache-2.0
Devopsnicepkg/auto-company1922 repos~814Automated safety check: PassMIT
Kcli Cluster Deploymentkarmab/kcli653—~1.5kAutomated safety check: PassApache-2.0
Google Agents CLI Publishpifferologo/cloud-agents-cli1291 repos~2.4kAutomated safety check: PassApache-2.0
DeployingGoogleCloudPlatform/race-condition234—~3kAutomated safety check: PassCustom licence
Aicr Uat ReportNVIDIA/aicr439—~3.2kAutomated safety check: PassApache-2.0

Similar skills

  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    192 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • Guides deployment and management of Kubernetes clusters with kcli.

    653 GitHub stars~1.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Google Agents CLI Publish

    pifferologo/cloud-agents-cli

    This skill should be used when the user wants to "publish an agent", "publish my ADK agent", "register an agent with Gemini Enterprise", "publish to Gemini Enterprise", or needs guidance on the…

    129 GitHub starsUsed in 1 repo~2.4k tokens
    DevOps & CloudAuto-check passed
  • Deploying

    GoogleCloudPlatform/race-condition

    Guides deployment of Race Condition to a GCP project. An agent skill from GoogleCloudPlatform/race-condition.

    234 GitHub stars~3k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Aicr Uat Report

    NVIDIA/aicr

    Official

    A skill your agent uses when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run…

    439 GitHub stars~3.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • GCP Secret Manager

    sickn33/agentic-awesome-skills

    Secure secrets in Google Cloud Secret Manager. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 3 repos~3.3k tokens
    DevOps & CloudAuto-check passed

More from GoogleCloudPlatform/accelerated-platforms

  • LLM D Workload Tuner

    GoogleCloudPlatform/accelerated-platforms

    This is an experimental Skill. An agent skill from GoogleCloudPlatform/accelerated-platforms.

    106 GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed
  • LLM D Deploy Stack

    GoogleCloudPlatform/accelerated-platforms

    Deploys the llm-d stack on GKE using well-lit paths specification.

    106 GitHub stars~3k tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about Gke Inference Stack Deploy

What does Gke Inference Stack Deploy do?

Deploys the GKE base platform and inference-specific terra-services (GPU/TPU) for accelerated workloads. Gke Inference Stack Deploy is an agent skill from GoogleCloudPlatform/accelerated-platforms. Deploys the GKE base platform and inference-specific terra-services (GPU/TPU) for accelerated workloads.

When should I use Gke Inference Stack Deploy?

Gke Inference Stack Deploy fits situations like: devOps & Cloud work in your project.

How do I install Gke Inference Stack Deploy in Claude Code?

Run `npx skills add GoogleCloudPlatform/accelerated-platforms --skill gke-inference-stack-deploy -a claude-code`. Or copy the skill folder (skills/gke-inference-stack-deploy in GoogleCloudPlatform/accelerated-platforms) into .claude/skills/gke-inference-stack-deploy in your project. Claude Code loads it when a task matches its description.

How do I install Gke Inference Stack Deploy in Codex?

Run `npx skills add GoogleCloudPlatform/accelerated-platforms --skill gke-inference-stack-deploy -a codex`. Or copy the skill folder (skills/gke-inference-stack-deploy in GoogleCloudPlatform/accelerated-platforms) into .agents/skills/gke-inference-stack-deploy in your project. Codex loads it when a task matches its description.

Can I use Gke Inference Stack Deploy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GoogleCloudPlatform/accelerated-platforms --skill gke-inference-stack-deploy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-inference-stack-deploy, .gemini/skills/gke-inference-stack-deploy, .github/skills/gke-inference-stack-deploy and .opencode/skills/gke-inference-stack-deploy in your project.

What does Gke Inference Stack Deploy need to run?

Going by SKILL.md and its folder, Gke Inference Stack Deploy needs the command-line tools its instructions call (terraform, kubectl and gcloud). Our summary lists: Python 3. Its frontmatter pre-approves these tools: kubectl, gcloud, terraform, python3.

Does Gke Inference Stack Deploy access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gke Inference Stack Deploy safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gke Inference Stack Deploy use?

Gke Inference Stack Deploy is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gke Inference Stack Deploy use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gke Inference Stack Deploy?

Skills that share tags, products or a category with Gke Inference Stack Deploy: Devops (nicepkg/auto-company, 192 stars), Kcli Cluster Deployment (karmab/kcli, 653 stars), Google Agents CLI Publish (pifferologo/cloud-agents-cli, 129 stars) and Deploying (GoogleCloudPlatform/race-condition, 234 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gke Inference Stack Deploy?

GoogleCloudPlatform (a GitHub organization) maintains it in GoogleCloudPlatform/accelerated-platforms, which has 106 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 5, 2026.

Source: GoogleCloudPlatform/accelerated-platforms on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.