Deploys the llm-d stack on GKE using well-lit paths specification.

Apache-2.0Auto-check passedDevOps & Cloud

Install LLM D Deploy Stack

skills CLI
$ npx skills add GoogleCloudPlatform/accelerated-platforms --skill llm-d-deploy-stack -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GoogleCloudPlatform/accelerated-platforms llm-d-deploy-stack --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GoogleCloudPlatform/accelerated-platforms.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-d-deploy-stack .claude/skills/llm-d-deploy-stack && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-d-deploy-stack
GitHub stars
106
Token cost
~3k tokens
SKILL.md length
1,011 words
Files
2
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deploys the llm-d stack on GKE using well-lit paths specification.

  • Works in 6 steps: Prerequisites & Cluster Setup → Configure Model & Accelerator → Select & Deploy Well-Lit Path Guide → …
  • DevOps & Cloud work in your project
  • SKILL.md covers Terminology & Variables, 1. Prerequisites & Cluster Setup, 2. Configure Model & Accelerator and 3. Select & Deploy Well-Lit…, plus 3 more sections
  • Calls kubectl and gcloud; needs HF_TOKEN

What it does

LLM D Deploy Stack is an agent skill from GoogleCloudPlatform/accelerated-platforms. Deploys the llm-d stack on GKE using well-lit paths specification.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/evals.json`).

It sits in DevOps & Cloud. It works with Google Kubernetes Engine and Google Cloud. The repository describes itself as: This repository is a collection of accelerated platform best practices, reference architectures, example use cases, reference implementations, and various other assets on Google… The licence is Apache-2.0.

When your agent uses it

  • DevOps & Cloud work in your project

Example prompts

  • “Use the llm-d-deploy-stack skill to deploy the llm-d stack on GKE using well-lit paths specification”
  • “/llm-d-deploy-stack”

Requirements

  • Python 3
  • A credential in YOUR_HUGGINGFACE_READ_TOKEN
  • Pre-approved tools (allowed-tools): kubectl, gcloud, helm, kustomize, curl, terraform, python3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Prerequisites & Cluster Setup
  2. Configure Model & Accelerator
  3. Select & Deploy Well-Lit Path Guide
  4. Hugging Face Token Setup
  5. Deploy Model Download Job & Model Server
  6. Verification

What it can do on your machine

Read from SKILL.md and the folder at commit bab190a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • kubectl
    • gcloud
    • helm
    • kustomize
    • curl
    • terraform
    • python3

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • gcloud

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl and gcloud, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM D Deploy Stack loads about 3k tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 1,011 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GoogleCloudPlatform/accelerated-platforms at commit bab190a, republished under its Apache-2.0 licence (© GoogleCloudPlatform). 1,011 words, ~3,042 tokens.

Download SKILL.mdSave it as .claude/skills/llm-d-deploy-stack/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
llm-d-deploy-stack
description
Deploys the llm-d stack on GKE using well-lit paths specification.
allowed-tools
kubectl, gcloud, helm, kustomize, curl, terraform, python3
version
0.4.0

Deploy llm-d Stack Skill

Follow these instructions to deploy the llm-d benchmarking stack on GKE.

Terminology & Variables

This skill utilizes the following core variables:

  • SPEC (also referred to as strategy or guide): This represents the name of the llm-d well-lit path guide being targeted. It maps directly to GKE overlay folder structures (llmd-<spec>)

    • Allowed Values:
      • optimized-baseline (Standard optimized baseline configuration).
      • precise-prefix-cache-routing (Includes precise cache routing for multi-turn workloads).
      • predicted-latency-routing (Includes dynamic latency-based routing overlays).
      • pd-disaggregation (Separates prefill and decode onto dedicated model servers. TPU v6e only, and only with the qwen/qwen3-32b model).
  • CRITICAL WARNING:

    1. NEVER run any teardown-*.sh scripts (e.g., teardown-llmd-optimized-baseline.sh) when attempting to "fix and rerun" a deployment or remove a workload. These scripts default to a full core platform teardown (ACP_TEARDOWN_CORE_PLATFORM=true) and will completely destroy the GKE cluster and all its resources. If a deployment fails, debug in place or re-run the deploy scripts. ONLY run teardown script when the user explicitly asks to tear down the stack and confirm with you first.
    2. NEVER Store a secret value or retrieve a value of a secret in memory or write it to any file. Instead, always use gcloud secrets describe [SECRET_NAME] OR kubectl get kubectl get secret [SECRET_NAME] -o yaml. If a secret does not exist ask the user to create it.

1. Prerequisites & Cluster Setup

  1. Ask the user: "Is there an existing GKE cluster? (yes/no)"
  2. If NO:
    • Ask the user: "What platform name would you like to use for the new cluster, and what is your Google Cloud project ID?"
    • Use sed to inject these values into ${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars:
      bash
      sed -i 's/^platform_name.*/platform_name = "<platform_name>"/g' "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars"
      # If platform_default_project_id doesn't exist, append it:
      grep -q "^platform_default_project_id" "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars" || echo "platform_default_project_id = \"\"" >> "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars"
      sed -i 's/^platform_default_project_id.*/platform_default_project_id = "<project_id>"/g' "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars"
    • Proceed to Section 2: Configure Model & Accelerator (to configure variables and create cluster).
  3. If YES:
    • Verify connectivity to the target cluster:
      bash
      kubectl cluster-info
    • Ask the user: "Is the llm-d stack already deployed on this cluster? (yes/no)"
    • If YES:
      • Skip to Section 4: Hugging Face Token Setup (we only need to deploy the model server, which is done in Section 5).
    • If NO:
      • Proceed to Section 2: Configure Model & Accelerator (to configure the deployment variables).

2. Configure Model & Accelerator

  1. Ask the user: "Which model would you like to run?"
    • Allowed Models:
      • google/gemma-4-31b-it
      • qwen/qwen3-32b (default)
      • qwen/qwen3-32b-fp8
      • redhatai/gemma-4-31b-it-fp8-block
    • Validation: If the user inputs anything else, flag it as invalid, display the allowed models, and ask again.
  2. Ask the user: "Which accelerator would you like to use?"
    • Allowed Accelerators:
      • rtx-pro-6000 (default)
      • h100 (translates to nvidia-h100)
      • h200 (translates to nvidia-h200)
      • v6e (TPU, translates to google-tpu-v6e)
    • Validation: If the user inputs anything else, flag it as invalid, display the allowed accelerators, and ask again.
  3. Update configuration:
    • Use sed to inject the chosen model and accelerator into ${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/llmd.auto.tfvars:
      bash
      echo "llmd_model_id = \"<chosen_model>\"" >> "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/llmd-shared.auto.tfvars"
      echo "llmd_accelerator_type = \"<chosen_accelerator>\"" >> "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/llmd-shared.auto.tfvars"

3. Select & Deploy Well-Lit Path Guide

  1. Ask the user: "Which llm-d well-lit path guide would you like to deploy?"

  2. Deploy the baseline stack: Run the deployment script corresponding to the chosen guide to create the cluster (if new) and deploy the baseline infra/services:

    • For optimized-baseline:
      bash
      "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/deploy-llmd-optimized-baseline.sh"
    • For precise-prefix-cache-routing:
      bash
      "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/deploy-llmd-precise-prefix-cache-routing.sh"
    • For predicted-latency-routing:
      bash
      "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/deploy-llmd-predicted-latency-routing.sh"
    • For pd-disaggregation:
      bash
      "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/deploy-llmd-pd-disaggregation.sh"
  3. Run validation:

    • Run the following command to verify the chosen custom compute class exists on the cluster:
      bash
      kubectl get computeclasses
    • Ensure that the accelerator type you chose appears in the list. If it does not, warn the user that the accelerator is unsupported or missing its custom compute class.
    • For pd-disaggregation, the required compute class is tpu-v6e-2x4 (8 chips), not the tpu-v6e-2x2 used by the other guides.
Show full SKILL.md (373 more words)Show less

4. Hugging Face Token Setup

  1. Instruct the user to add their Hugging Face Read Token to Google Secret Manager and as a Kubernetes secret: Provide them with these commands, replacing <YOUR_HUGGINGFACE_READ_TOKEN> with their actual token. Note that the source command must be run in the same shell session as the subsequent commands so the environment variables are preserved:

    bash
    # Source environment variables
    source "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/scripts/set_environment_variables.sh"
    
    # Add to Secret Manager
    HF_TOKEN_READ=<YOUR_HUGGINGFACE_READ_TOKEN>
    echo ${HF_TOKEN_READ} | gcloud secrets versions add ${huggingface_hub_access_token_read_secret_manager_secret_name} --data-file=- --project=${huggingface_secret_manager_project_id}
    
    # Add to Kubernetes
    kubectl -n ${llmd_namespace} create secret generic llm-d-hf-token --from-literal=HF_TOKEN="${HF_TOKEN_READ}"
  2. WAIT: Stop calling tools and ask the user to confirm once they add the HF token to secret manager and kubernetes secret Do not proceed until the user confirms.

5. Deploy Model Download Job & Model Server

Once the user confirms the token is configured, proceed with the deployment:

  1. Deploy the model download job:
    bash
    # Configure
    "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/kubernetes-manifests/model-download/configure_huggingface.sh"
    # Apply
    kubectl apply --kustomize "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/kubernetes-manifests/model-download/huggingface"
  2. Wait for download to complete: Monitor the job:
    bash
    kubectl get job -n ${huggingface_hub_downloader_kubernetes_namespace_name}
    Wait until the job status shows Complete.
  3. Clean up the download job: Once the download is complete, delete the job to free up GKE resources:
    bash
    kubectl delete job -n ${huggingface_hub_downloader_kubernetes_namespace_name} ${HF_MODEL_ID_HASH}-hf-model-to-gcs
  4. Deploy the Model Server:
    • Configure the model server (run the script for GPU or TPU as indicated by the chosen accelerator):
      • If GPU:
        bash
        "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/kubernetes-manifests/online-inference-gpu/llmd-<spec>/vllm/configure_vllm.sh"
      • If TPU:
        bash
        "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/kubernetes-manifests/online-inference-tpu/llmd-<spec>/vllm/configure_vllm.sh"
    • Deploy using the appropriate overlay directory. Construct the directory path as: platforms/gke/base/use-cases/inference-ref-arch/kubernetes-manifests/online-inference-[gpu|tpu]/llmd-[spec]/vllm/[prefix]-[suffix]
      • [gpu|tpu]: Use tpu if the accelerator is v6e, otherwise gpu.
      • [spec]: The well-lit path chosen in Section 3.
      • [prefix]: The accelerator prefix (e.g., rtx-pro-6000, h100, h200, v6e).
      • [suffix]: The model name suffix (e.g., gemma-4-31b-it, qwen3-32b).
      bash
      kubectl apply --kustomize "${ACP_REPO_DIR}/<constructed_overlay_dir>"

6. Verification

  • Check that all pods are running and services are accessible. (Note: You may need to source the environment variables script first to get $llmd_namespace):
    bash
    source "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/scripts/set_environment_variables.sh"
    kubectl get pods -n ${llmd_namespace}
    kubectl get svc -n ${llmd_namespace}
  • Verify the model downloader job has completed and the model files are in the GCS bucket:
    bash
    source "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/scripts/set_environment_variables.sh"
    kubectl get jobs -n ${llmd_namespace}
    gcloud storage ls gs://${huggingface_hub_models_bucket_name}/${llmd_model_id}/
  • Ensure the Hugging Face token is securely configured:
    bash
    source "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/scripts/set_environment_variables.sh"
    kubectl describe secretProviderClass huggingface-tokens -n ${llmd_namespace}
    kubectl describe secret llm-d-hf-token -n ${llmd_namespace}

© GoogleCloudPlatform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/llm-d-deploy-stack of GoogleCloudPlatform/accelerated-platforms.

  • SKILL.md
  • evals/evals.json

Open the folder on GitHubat commit bab190a

Compare with similar skills

LLM D Deploy Stack next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM D Deploy Stack compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM D Deploy Stack this skillGoogleCloudPlatform/accelerated-platforms106—~3kAutomated safety check: PassApache-2.0
Devopsnicepkg/auto-company1922 repos~814Automated safety check: PassMIT
Kcli Cluster Deploymentkarmab/kcli653—~1.5kAutomated safety check: PassApache-2.0
Google Agents CLI Publishpifferologo/cloud-agents-cli1291 repos~2.4kAutomated safety check: PassApache-2.0
DeployingGoogleCloudPlatform/race-condition234—~3kAutomated safety check: PassCustom licence
Aicr Uat ReportNVIDIA/aicr439—~3.2kAutomated safety check: PassApache-2.0

Similar skills

  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    192 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • Guides deployment and management of Kubernetes clusters with kcli.

    653 GitHub stars~1.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Google Agents CLI Publish

    pifferologo/cloud-agents-cli

    This skill should be used when the user wants to "publish an agent", "publish my ADK agent", "register an agent with Gemini Enterprise", "publish to Gemini Enterprise", or needs guidance on the…

    129 GitHub starsUsed in 1 repo~2.4k tokens
    DevOps & CloudAuto-check passed
  • Deploying

    GoogleCloudPlatform/race-condition

    Guides deployment of Race Condition to a GCP project. An agent skill from GoogleCloudPlatform/race-condition.

    234 GitHub stars~3k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Aicr Uat Report

    NVIDIA/aicr

    Official

    A skill your agent uses when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run…

    439 GitHub stars~3.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • GCP Secret Manager

    sickn33/agentic-awesome-skills

    Secure secrets in Google Cloud Secret Manager. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 3 repos~3.3k tokens
    DevOps & CloudAuto-check passed

More from GoogleCloudPlatform/accelerated-platforms

  • LLM D Workload Tuner

    GoogleCloudPlatform/accelerated-platforms

    This is an experimental Skill. An agent skill from GoogleCloudPlatform/accelerated-platforms.

    106 GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed
  • Gke Inference Stack Deploy

    GoogleCloudPlatform/accelerated-platforms

    Deploys the GKE base platform and inference-specific terra-services (GPU/TPU) for accelerated workloads.

    106 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about LLM D Deploy Stack

What does LLM D Deploy Stack do?

Deploys the llm-d stack on GKE using well-lit paths specification. LLM D Deploy Stack is an agent skill from GoogleCloudPlatform/accelerated-platforms. Deploys the llm-d stack on GKE using well-lit paths specification.

When should I use LLM D Deploy Stack?

LLM D Deploy Stack fits situations like: devOps & Cloud work in your project.

How do I install LLM D Deploy Stack in Claude Code?

Run `npx skills add GoogleCloudPlatform/accelerated-platforms --skill llm-d-deploy-stack -a claude-code`. Or copy the skill folder (skills/llm-d-deploy-stack in GoogleCloudPlatform/accelerated-platforms) into .claude/skills/llm-d-deploy-stack in your project. Claude Code loads it when a task matches its description.

How do I install LLM D Deploy Stack in Codex?

Run `npx skills add GoogleCloudPlatform/accelerated-platforms --skill llm-d-deploy-stack -a codex`. Or copy the skill folder (skills/llm-d-deploy-stack in GoogleCloudPlatform/accelerated-platforms) into .agents/skills/llm-d-deploy-stack in your project. Codex loads it when a task matches its description.

Can I use LLM D Deploy Stack in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GoogleCloudPlatform/accelerated-platforms --skill llm-d-deploy-stack -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-d-deploy-stack, .gemini/skills/llm-d-deploy-stack, .github/skills/llm-d-deploy-stack and .opencode/skills/llm-d-deploy-stack in your project.

What does LLM D Deploy Stack need to run?

Going by SKILL.md and its folder, LLM D Deploy Stack needs the command-line tools its instructions call (kubectl and gcloud) and credentials named HF_TOKEN. Our summary lists: Python 3; A credential in YOUR_HUGGINGFACE_READ_TOKEN. Its frontmatter pre-approves these tools: kubectl, gcloud, helm, kustomize, curl, terraform, python3.

Does LLM D Deploy Stack access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is LLM D Deploy Stack safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM D Deploy Stack use?

LLM D Deploy Stack is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM D Deploy Stack use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM D Deploy Stack?

Skills that share tags, products or a category with LLM D Deploy Stack: Devops (nicepkg/auto-company, 192 stars), Kcli Cluster Deployment (karmab/kcli, 653 stars), Google Agents CLI Publish (pifferologo/cloud-agents-cli, 129 stars) and Deploying (GoogleCloudPlatform/race-condition, 234 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM D Deploy Stack?

GoogleCloudPlatform (a GitHub organization) maintains it in GoogleCloudPlatform/accelerated-platforms, which has 106 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 5, 2026.

Source: GoogleCloudPlatform/accelerated-platforms on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.