Official agent skill

Gke Workload Scaling

by google in google/skills

Manages scaling for GKE workloads using HPA and VPA. An agent skill from google/skills.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Gke Workload Scaling

skills CLI
$ npx skills add google/skills --skill gke-workload-scaling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills gke-workload-scaling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-workload-scaling .claude/skills/gke-workload-scaling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gke-workload-scaling
GitHub stars
21k
Token cost
~1.3k tokens
SKILL.md length
516 words
Files
3 (incl. assets)
Skills in repo
147
Repo updated
First seen
Licence
Apache-2.0

At a glance

Manages scaling for GKE workloads using HPA and VPA. An agent skill from google/skills.

  • Works in 3 steps: Manual Scaling → Horizontal Pod Autoscaling (HPA) → Vertical Pod Autoscaling (VPA)
  • Configuring Horizontal Pod Autoscaler (HPA)
  • SKILL.md covers Workflows, Best Practices and Rightsizing Workflow
  • Calls kubectl and gcloud

What it does

Gke Workload Scaling is an agent skill from google/skills, published by the product's own GitHub organization. Manages scaling for GKE workloads using HPA and VPA. Use when configuring Horizontal Pod Autoscaler (HPA), configuring Vertical Pod Autoscaler (VPA), or applying best practices for GKE workload autoscaling. Do not use for cluster-level autoscaling (Cluster Autoscaler), static cluster sizing, or configuring node-level machine styles directly.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including assets (for example `assets/hpa-example.yaml` and `assets/vpa-example.yaml`).

It sits in DevOps & Cloud. It works with Google Kubernetes Engine and Prometheus. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • Configuring Horizontal Pod Autoscaler (HPA)
  • Configuring Vertical Pod Autoscaler (VPA)
  • Applying best practices for GKE workload autoscaling
  • Cluster-level autoscaling (Cluster Autoscaler)

Example prompts

  • “Use the gke-workload-scaling skill to manage scaling for GKE workloads using HPA and VPA. An agent skill from google/skills”
  • “/gke-workload-scaling”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Manual Scaling
  2. Horizontal Pod Autoscaling (HPA)
  3. Vertical Pod Autoscaling (VPA)

What it can do on your machine

Read from SKILL.md and the folder at commit 7d97937. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • gcloud

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl and gcloud, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gke Workload Scaling loads about 1.3k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 516 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from google/skills at commit 7d97937, republished under its Apache-2.0 licence (© google). 516 words, ~1,254 tokens.

Download SKILL.mdSave it as .claude/skills/gke-workload-scaling/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
gke-workload-scaling
description
Manages scaling for GKE workloads using HPA and VPA. Use when configuring Horizontal Pod Autoscaler (HPA), configuring Vertical Pod Autoscaler (VPA), or applying best practices for GKE workload autoscaling. Do not use for cluster-level autoscaling (Cluster Autoscaler), static cluster sizing, or configuring node-level machine styles directly.
metadata.version
1.0.0
metadata.category
Containers

GKE Workload Scaling

This skill provides workflows and best practices for scaling applications on Google Kubernetes Engine (GKE). It covers manual scaling, Horizontal Pod Autoscaling (HPA), and Vertical Pod Autoscaling (VPA).

Workflows

1. Manual Scaling

Scale a deployment to a fixed number of replicas. Useful for immediate manual intervention or testing.

Command:

bash
kubectl scale deployment {deployment_name} --replicas={number} -n {namespace}

# Verify the scale event
kubectl get deployment {deployment_name} -n {namespace}
2. Horizontal Pod Autoscaling (HPA)

Automatically scale the number of pods based on observed CPU utilization, memory utilization, or custom metrics.

Prerequisites:

  • Metrics Server must be running (enabled by default on GKE).
  • Containers clearly define resource requests/limits.

Quick Command:

bash
kubectl autoscale deployment {deployment_name} --cpu-percent=50 --min=1 --max=10

Manifest Approach (Recommended): Use a YAML manifest for version-controlled configuration. See assets/hpa-example.yaml for a template.

bash
kubectl apply -f assets/hpa-example.yaml

# Verify HPA is created and fetching metrics
kubectl get hpa

Custom Metrics & External Metrics: For GKE, the modern and recommended approach for scaling based on Cloud Monitoring metrics (e.g., Pub/Sub queue length) is to use the External metric type, which is natively supported by the GKE control plane without requiring the Custom Metrics Adapter. For application-specific metrics exposed via Prometheus, you can use Google Cloud Managed Service for Prometheus or the Prometheus Adapter.

3. Vertical Pod Autoscaling (VPA)

Automatically adjust the CPU and memory reservations for your pods to match actual usage. This is critical for right-sizing workloads.

Prerequisites:

  • VPA must be enabled on the cluster.
    • Autopilot: Enabled by default.
    • Standard: Must be enabled manually.

Enable VPA on Standard Cluster:

bash
gcloud container clusters update {cluster_name} --enable-vertical-pod-autoscaling --zone {zone}

Update Modes:

  • Off: Calculates recommendations but does not apply them. Good for "dry run" analysis.
  • Initial: Assigns resources only at pod creation time.
  • Auto: Updates running pods by restarting them if recommendations differ significantly from requests.
  • InPlaceOrRecreate: Attempts to update Pod resources without recreating the Pod. If in-place update is not possible, it reverts to Auto mode (requires GKE 1.34+).

Example: See assets/vpa-example.yaml for a configuration template.

Show full SKILL.md (233 more words)Show less

Best Practices

  1. Define Resource Requests: HPA and VPA rely on accurate resource requests. Always define them in your container specs.
  2. Avoid Metric Conflicts: Do not configure HPA and VPA to use the same metric (e.g., both CPU). This causes thrashing.
    • Typical Pattern: HPA on CPU, VPA on Memory.
  3. Pod Disruption Budgets (PDBs): Define PDBs to ensure application availability during scaling events or node upgrades.
  4. HPA Lag: HPA has a stabilization window (default 5 mins) to prevent rapid fluctuation.
  5. VPA "Auto" Mode Risks: In "Auto" mode, VPA restarts pods to change resources. Ensure your application handles restarts gracefully (e.g., handles SIGTERM).
    • Note: By default, VPA requires at least 2 replicas to perform evictions (to prevent a situation where the only running replica is evicted, causing downtime). In GKE 1.22+, you can override this by setting minReplicas in PodUpdatePolicy.

Rightsizing Workflow

  1. Deploy VPA in Off mode for 24+ hours
  2. Read recommendations: kubectl describe vpa {deployment_name}-vpa -n {namespace}
  3. Compare target values against current requests
  4. Apply with 20% buffer: new_request = target * 1.2
  5. Use patch format or update deployment manifest to apply new resource requests
ConditionRecommendationRisk
CPU request >5x P95 actualReduce to P95 * 1.2Medium
Memory request >3x P95 actualReduce to P95 * 1.2Medium
CPU request >2x P95 actualRightsizing with 20% bufferLow
No resource limits setAdd limits to prevent noisy-neighborLow

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (assets) in skills/cloud/gke-workload-scaling of google/skills.

  • SKILL.md
  • assets/hpa-example.yaml
  • assets/vpa-example.yaml

Open the folder on GitHubat commit 7d97937

Compare with similar skills

Gke Workload Scaling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gke Workload Scaling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gke Workload Scaling this skillgoogle/skills21k—~1.3kAutomated safety check: PassApache-2.0
Syncmetapawurb/hotpath-rs1.9k—~1.2kAutomated safety check: NotesMIT
Devopsnicepkg/auto-company1922 repos~814Automated safety check: PassMIT
KubeShark for KubernetesLukasNiessen/kubernetes-skill444—~1.2kAutomated safety check: PassMIT
Optimize Slurm TopologyNVlabs/alpasim1.3k—~1.6kAutomated safety check: PassApache-2.0
UModel Root Cause Analysisalibaba/UnifiedModel412—~1.9kAutomated safety check: PassCustom licence

Similar skills

  • Syncmeta

    pawurb/hotpath-rs

    Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta).

    1.9k GitHub stars~1.2k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    192 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • KubeShark for Kubernetes

    LukasNiessen/kubernetes-skill

    Keeps Kubernetes manifests, Helm charts and policies grounded by diagnosing six failure modes, such as insecure defaults and API drift, and loading only matching references.

    444 GitHub stars~1.2k tokensUpdated 25 days ago
    DevOps & CloudAuto-check passed
  • Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.

    1.3k GitHub stars~1.6k tokensUpdated 20 days ago
    DevOps & CloudAuto-check passed
  • UModel Root Cause Analysis

    alibaba/UnifiedModel

    Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.

    412 GitHub stars~1.9k tokensUpdated 15 days ago
    DevOps & CloudAuto-check passed
  • Greptimedb Perses Dashboard

    GreptimeTeam/dashboard

    Generate Perses dashboards or single panels for GreptimeDB. An agent skill from GreptimeTeam/dashboard.

    111 GitHub stars~3.9k tokensUpdated 8 days ago
    DevOps & CloudAuto-check passed

More from google/skills

All 147 skills in this repo
  • Official

    Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.

    21k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated yesterday
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Gke Workload Scaling

What does Gke Workload Scaling do?

Manages scaling for GKE workloads using HPA and VPA. An agent skill from google/skills. Gke Workload Scaling is an agent skill from google/skills, published by the product's own GitHub organization. Manages scaling for GKE workloads using HPA and VPA.

When should I use Gke Workload Scaling?

Gke Workload Scaling fits situations like: configuring Horizontal Pod Autoscaler (HPA); configuring Vertical Pod Autoscaler (VPA); applying best practices for GKE workload autoscaling; cluster-level autoscaling (Cluster Autoscaler).

How do I install Gke Workload Scaling in Claude Code?

Run `npx skills add google/skills --skill gke-workload-scaling -a claude-code`. Or copy the skill folder (skills/cloud/gke-workload-scaling in google/skills) into .claude/skills/gke-workload-scaling in your project. Claude Code loads it when a task matches its description.

How do I install Gke Workload Scaling in Codex?

Run `npx skills add google/skills --skill gke-workload-scaling -a codex`. Or copy the skill folder (skills/cloud/gke-workload-scaling in google/skills) into .agents/skills/gke-workload-scaling in your project. Codex loads it when a task matches its description.

Can I use Gke Workload Scaling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-workload-scaling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-workload-scaling, .gemini/skills/gke-workload-scaling, .github/skills/gke-workload-scaling and .opencode/skills/gke-workload-scaling in your project.

What does Gke Workload Scaling need to run?

Going by SKILL.md and its folder, Gke Workload Scaling needs the command-line tools its instructions call (kubectl and gcloud).

Does Gke Workload Scaling access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gke Workload Scaling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gke Workload Scaling use?

Gke Workload Scaling is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gke Workload Scaling use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gke Workload Scaling?

Skills that share tags, products or a category with Gke Workload Scaling: Syncmeta (pawurb/hotpath-rs, 1.9k stars), Devops (nicepkg/auto-company, 192 stars), KubeShark for Kubernetes (LukasNiessen/kubernetes-skill, 444 stars) and Optimize Slurm Topology (NVlabs/alpasim, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gke Workload Scaling?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,032 GitHub stars. The repository holds 147 skills in this directory. The repository was last updated on October 8, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.