This is an experimental Skill. An agent skill from GoogleCloudPlatform/accelerated-platforms.

Apache-2.0Auto-check passedAI & LLM Engineering

Install LLM D Workload Tuner

skills CLI
$ npx skills add GoogleCloudPlatform/accelerated-platforms --skill llm-d-workload-tuner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GoogleCloudPlatform/accelerated-platforms llm-d-workload-tuner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GoogleCloudPlatform/accelerated-platforms.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-d-workload-tuner .claude/skills/llm-d-workload-tuner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-d-workload-tuner
GitHub stars
106
Token cost
~1.2k tokens
SKILL.md length
424 words
Files
5 (incl. scripts, references)
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

This is an experimental Skill. An agent skill from GoogleCloudPlatform/accelerated-platforms.

  • Works in 5 steps: Terminology & Variables → Prerequisites → Running the Tuner → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers 1. Terminology & Variables, 2. Prerequisites, 3. Running the Tuner and 4. Applying the Tuned…, plus 1 more section
  • Runs Python scripts from its folder; calls python3 and kubectl

What it does

LLM D Workload Tuner is an agent skill from GoogleCloudPlatform/accelerated-platforms. This is an experimental Skill. It automatically tunes GKE vLLM inference server parameters and resources based on workload profiles specified in the benchmark configs.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `evals/evals.json`, `references/llm-d-workload-profiles.md` and `references/model_specs.json`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Google Kubernetes Engine and vLLM. The repository describes itself as: This repository is a collection of accelerated platform best practices, reference architectures, example use cases, reference implementations, and various other assets on Google… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/llm-d-workload-tuner”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): python3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Terminology & Variables
  2. Prerequisites
  3. Running the Tuner
  4. Applying the Tuned Configuration
  5. Verification

What it can do on your machine

Read from SKILL.md and the folder at commit bab190a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • python3

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM D Workload Tuner loads about 1.2k tokens when it runs, and up to ~1.9k if it reads all its reference files. Until then it costs about 47 tokens; SKILL.md has 424 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~47
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from GoogleCloudPlatform/accelerated-platforms at commit bab190a, republished under its Apache-2.0 licence (© GoogleCloudPlatform). 424 words, ~1,209 tokens.

Download SKILL.mdSave it as .claude/skills/llm-d-workload-tuner/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
llm-d-workload-tuner
description
This is an experimental Skill. It automatically tunes GKE vLLM inference server parameters and resources based on workload profiles specified in the benchmark configs.
allowed-tools
python3
version
0.1.0

LLM-D Workload Tuner Skill

Follow these instructions to run the workload optimizer and optionally apply the tuned vLLM parameters and GKE node resource requirements.

1. Terminology & Variables

This skill utilizes the following core variables:

  • SPEC (passed via --spec flag): The target GKE routing specification/strategy overlay name (e.g., optimized-baseline, precise-prefix-cache-routing, predicted-latency-routing). It defines which subdirectory under GKE manifests will be read and patched.
  • Workload Profile: The benchmark workload specification file (e.g., chatbot_synthetic.yaml.in, agentic_code_generation.yaml.in) which defines the load-testing prompt distributions and sequence limits.

2. Prerequisites

  • Review the target workload characteristics and configurations in the reference guide: llm-d-workload-profiles.md.
  • Check model specifications in model_specs.json, which hosts the 6-element tuple defining hardware parameters and the maximum supported context length ceiling ([Parameters (B), Layers, KV Heads, Head Dimension, Suffix, Max Context Length]).
  • Ensure you know the path to your llm-d-benchmark directory. The benchmarking files config.json (defining sequence lengths) and inference-perf.yaml (defining stages and model servers) come from the specific llm-d-benchmark workload profile chosen and must exist under the target benchmarking directories.
  • Know the target accelerator type (e.g., rtx-pro-6000, nvidia-h100, or v6e TPU)
  • Deployment strategy spec name (e.g., precise-prefix-cache-routing, optimized-baseline, or predicted-latency-routing). You must explicitly supply the target spec name using the --spec flag.

3. Running the Tuner

To compute optimal sizing (Tensor Parallelism size, maximum model length limits, memory margins, and chunked prefill). Note that the --spec parameter is required:

  • Command Format:

    bash
    python3 "${ACP_REPO_DIR}/skills/llm-d-workload-tuner/scripts/tune_workload.py" \
        [--config <path_to_config.json>] \
        --perf-yaml <path_to_inference-perf.yaml> \
        --accelerator-type <accelerator_name> \
        --spec <spec_name>
  • Example (Dry Run):

    bash
    python3 "${ACP_REPO_DIR}/skills/llm-d-workload-tuner/scripts/tune_workload.py" \
        --perf-yaml llm-d-benchmark/workload/profiles/inference-perf/chatbot_synthetic.yaml.in \
        --accelerator-type rtx-pro-6000 \
        --spec precise-prefix-cache-routing
Show full SKILL.md (221 more words)Show less

4. Applying the Tuned Configuration

Confirm the GKE cluster name, region and reservation (if ANY) and then use the --apply flag to commit the calculated tuning configs directly to the GKE deployment overlays.

  • Command:
    bash
    python3 "${ACP_REPO_DIR}/skills/llm-d-workload-tuner/scripts/tune_workload.py" \
        --perf-yaml llm-d-benchmark/workload/profiles/inference-perf/chatbot_synthetic.yaml.in \
        --accelerator-type rtx-pro-6000 \
        --spec precise-prefix-cache-routing \
        --apply

When --apply is set, the tuner:

  1. Updates runtime.env inside the GKE overlay directory (setting TENSOR_PARALLEL_SIZE and MAX_MODEL_LEN). KV cache sizing accounts for total sequence length (max_in + max_out), and MAX_MODEL_LEN is bounded by the architecture ceiling in model_specs.json.
  2. Patches patch-nodeselector.yaml to request matching GPU / TPU counts on nodes.
  3. Patches patch-resources.yaml to configure container GPU / TPU limit settings.
  4. Patches patch-tuner-args.yaml to configure optimal arguments for container index 0 (modelserver).
  5. Verifies and logs the parameter diff: Always review the printed === Configuration Gap Analysis === output to see exactly which parameters were tuned from the baseline deployment in this repo.

5. Verification

After deploying the tuned stack, verify:

  • That the baseline diff output matches the expected transitions.
  • That vLLM deployment specs match the calculated values:
    bash
    kubectl get deployment -n <namespace> -l app=vllm -o jsonpath='{.items[0].spec.template.spec.containers[0].args}'
  • Check if the user want to execute benchmark then call llm-d-benchmarking skill and pass the workload profile and endpoint url to it:
    bash
    "${ACP_REPO_DIR}/skills/llm-d-benchmarking/scripts/run_benchmark.sh" <workload_profile_name> <endpoint_url> [namespace] [model_name]
  • That benchmarking config are automatically bundled to GCS bucket along with the performance metrics

© GoogleCloudPlatform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/llm-d-workload-tuner of GoogleCloudPlatform/accelerated-platforms.

  • SKILL.md
  • evals/evals.json
  • references/llm-d-workload-profiles.md
  • references/model_specs.json
  • scripts/tune_workload.py

Open the folder on GitHubat commit bab190a

Compare with similar skills

LLM D Workload Tuner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM D Workload Tuner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM D Workload Tuner this skillGoogleCloudPlatform/accelerated-platforms106—~1.2kAutomated safety check: PassApache-2.0
Gke Manifest Generationgoogle/skills21k—~3.1kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Ascend Release Manager for vLLMvllm-project/vllm-ascend2.9k—~7.2kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters.

    21k GitHub stars~3.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Ascend Release Manager for vLLM

    vllm-project/vllm-ascend

    Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

    2.9k GitHub stars~7.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from GoogleCloudPlatform/accelerated-platforms

  • Gke Inference Stack Deploy

    GoogleCloudPlatform/accelerated-platforms

    Deploys the GKE base platform and inference-specific terra-services (GPU/TPU) for accelerated workloads.

    106 GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • LLM D Deploy Stack

    GoogleCloudPlatform/accelerated-platforms

    Deploys the llm-d stack on GKE using well-lit paths specification.

    106 GitHub stars~3k tokensUpdated 2 days ago
    Auto-check passed

Questions about LLM D Workload Tuner

What does LLM D Workload Tuner do?

This is an experimental Skill. An agent skill from GoogleCloudPlatform/accelerated-platforms. LLM D Workload Tuner is an agent skill from GoogleCloudPlatform/accelerated-platforms. This is an experimental Skill.

When should I use LLM D Workload Tuner?

LLM D Workload Tuner fits situations like: tasks that involve LLM inference and serving.

How do I install LLM D Workload Tuner in Claude Code?

Run `npx skills add GoogleCloudPlatform/accelerated-platforms --skill llm-d-workload-tuner -a claude-code`. Or copy the skill folder (skills/llm-d-workload-tuner in GoogleCloudPlatform/accelerated-platforms) into .claude/skills/llm-d-workload-tuner in your project. Claude Code loads it when a task matches its description.

How do I install LLM D Workload Tuner in Codex?

Run `npx skills add GoogleCloudPlatform/accelerated-platforms --skill llm-d-workload-tuner -a codex`. Or copy the skill folder (skills/llm-d-workload-tuner in GoogleCloudPlatform/accelerated-platforms) into .agents/skills/llm-d-workload-tuner in your project. Codex loads it when a task matches its description.

Can I use LLM D Workload Tuner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GoogleCloudPlatform/accelerated-platforms --skill llm-d-workload-tuner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-d-workload-tuner, .gemini/skills/llm-d-workload-tuner, .github/skills/llm-d-workload-tuner and .opencode/skills/llm-d-workload-tuner in your project.

What does LLM D Workload Tuner need to run?

Going by SKILL.md and its folder, LLM D Workload Tuner needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and kubectl). Our summary lists: Python 3. Its frontmatter pre-approves these tools: python3.

Does LLM D Workload Tuner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is LLM D Workload Tuner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does LLM D Workload Tuner use?

LLM D Workload Tuner is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM D Workload Tuner use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 656 tokens, read only when the agent opens those files.

What are the alternatives to LLM D Workload Tuner?

Skills that share tags, products or a category with LLM D Workload Tuner: Gke Manifest Generation (google/skills, 21k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Aider Delegate (amElnagdy/delegate-skills, 2.3k stars) and Hugging Face Local Model Evals (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM D Workload Tuner?

GoogleCloudPlatform (a GitHub organization) maintains it in GoogleCloudPlatform/accelerated-platforms, which has 106 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 5, 2026.

Source: GoogleCloudPlatform/accelerated-platforms on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.