Official agent skill

Gke Storage Troubleshooting

by google in google/skills

Diagnoses GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs, fsGroup mount timeouts), disk-performance and node storage-pressure issues, slow-disk…

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Gke Storage Troubleshooting

skills CLI
$ npx skills add google/skills --skill gke-storage-troubleshooting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills gke-storage-troubleshooting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-storage-troubleshooting .claude/skills/gke-storage-troubleshooting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gke-storage-troubleshooting
GitHub stars
21k
Token cost
~3.5k tokens
SKILL.md length
1,471 words
Files
1
Skills in repo
147
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnoses GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs, fsGroup mount timeouts), disk-performance and node storage-pressure issues, slow-disk…

  • Works in 4 steps: Non-Interactive Context Discovery &… → Classify the Storage Symptom → Resolution → …
  • Pods are stuck in ContainerCreating
  • SKILL.md covers 🔍 Diagnosis & Resolution… and References
  • Calls kubectl

What it does

Gke Storage Troubleshooting is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs, fsGroup mount timeouts), disk-performance and node storage-pressure issues, slow-disk Pod-creation failures, volume-expansion problems, Local SSD / Hyperdisk Storage Pool creation errors, and Cloud Storage FUSE OOM. Use when Pods are stuck in ContainerCreating, volumes fail to attach or mount, or nodes report storage pressure. Don't use for routine storage provisioning or StorageClass/PVC authoring (see the…

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud. It works with Google Kubernetes Engine and Google Cloud. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • Pods are stuck in ContainerCreating
  • Volumes fail to attach
  • Nodes report storage pressure
  • Routine storage provisioning

Example prompts

  • “Use the gke-storage-troubleshooting skill to diagnose GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs…”
  • “/gke-storage-troubleshooting”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Non-Interactive Context Discovery & Dry-Run Fallback
  2. Classify the Storage Symptom
  3. Resolution
  4. Propose the GitOps Correction

What it can do on your machine

Read from SKILL.md and the folder at commit 5120a76. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.cloud.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gke Storage Troubleshooting loads about 3.5k tokens when it runs. Until then it costs about 140 tokens; SKILL.md has 1,471 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~140
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from google/skills at commit 5120a76, republished under its Apache-2.0 licence (© google). 1,471 words, ~3,500 tokens.

Download SKILL.mdSave it as .claude/skills/gke-storage-troubleshooting/SKILL.md (or your agent's skills folder).
name
gke-storage-troubleshooting
description
Diagnoses GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs, fsGroup mount timeouts), disk-performance and node storage-pressure issues, slow-disk Pod-creation failures, volume-expansion problems, Local SSD / Hyperdisk Storage Pool creation errors, and Cloud Storage FUSE OOM. Use when Pods are stuck in ContainerCreating, volumes fail to attach or mount, or nodes report storage pressure. Don't use for routine storage provisioning or StorageClass/PVC authoring (see the gke-storage skill).
metadata.category
Storage
metadata.version
1.1.0

GKE Storage Troubleshooting Skill

Use this skill to systematically diagnose and resolve persistent-storage failures for workloads running on GKE — volume attach/mount errors, disk performance and node storage pressure, volume expansion, storage-related cluster/node-pool creation errors, and Cloud Storage FUSE memory issues. This skill operates non-interactively and enforces a read-only diagnostics boundary before proposing manifest or configuration corrections.

For routine storage provisioning and StorageClass/PVC authoring, use the gke-storage skill instead. This skill focuses on failure diagnosis.

🔍 Diagnosis & Resolution Workflow

Step 0: Non-Interactive Context Discovery & Dry-Run Fallback
  1. Parameter Extraction: Extract required context (project_id, cluster_name, cluster_location, workload_name, workload_namespace, pod_name, and the relevant pvc_name / pv_name / node_name) non-interactively from the user prompt, active SETTINGS.md, or environment defaults:

    • Default workload_namespace to default if omitted.
    • Infer missing cluster parameters from the active environment (kubectl config current-context or gcloud config get-value project).
  2. Cluster Credentials & Fallback Mode:

    • Attempt credential fetch: gcloud container clusters get-credentials {cluster_name} --location {cluster_location} --project {project_id}.
    • Fallback / Dry-Run Mode: If the cluster is unreachable, non-existent, or live command execution fails (such as in sandboxed evaluations, dry-run mode, or offline analysis):
      • Limit retry attempts to avoid resource exhaustion and context overflow.
      • Immediately present the exact kubectl / gcloud diagnostic commands for the human operator to run.
      • Synthesize the root-cause analysis and output the proposed GitOps correction based on the reported symptoms.

Step 1: Classify the Storage Symptom

Gather the primary signals, then jump to the matching branch under Step 2 (Resolution) — you normally perform only the one branch that matches your diagnosis, not all of them.

Diagnostic Commands:

bash
kubectl describe pod {pod_name} -n {workload_namespace}
kubectl get pvc,pv -n {workload_namespace}
kubectl get events -n {workload_namespace} --sort-by='.metadata.creationTimestamp'
kubectl describe node {node_name}
  • Pod stuck in ContainerCreating with an attach/mount event → Volume Attach & Mount Failures.
  • Node-level slowness, PLEG is not healthy, or StoragePressureDetected events → Disk Performance & Node Storage Pressure.
  • Cluster / node-pool creation or provisioning error → Storage Provisioning & Creation Failures.
  • A resized volume is not reflected inside the container → Volume Expansion Not Reflecting in the Container.
  • Cloud Storage FUSE Pod / sidecar OOM → Cloud Storage FUSE Out-Of-Memory (OOM) Events.

Step 2: Resolution

Perform only the branch that matches your Step 1 diagnosis. These branches are mutually exclusive alternatives, not sequential steps.

Volume Attach & Mount Failures
  • Error 400: Cannot attach RePD to an optimized VM: Regional persistent disks are restricted from being used with memory-optimized or compute-optimized machine types.

    • If a regional PD is not a hard requirement, switch the workload to a non-regional persistent disk StorageClass.
    • If a regional PD is required, use taints and tolerations so that Pods needing regional PDs are scheduled onto a node pool that does not use optimized machine types.
  • Pods stay Pending / FailedScheduling after a node pool is moved to a 4th-generation (N4, N4A, N4D) machine series while the workload uses a Persistent Disk StorageClass: N4/N4A/N4D machines do not support Persistent Disk (they support Hyperdisk only), so a PVC bound to a pd-* StorageClass cannot bind or schedule on those nodes. Events typically show FailedScheduling with a volume node-affinity / topology conflict.

    • Switch the workload to a Hyperdisk StorageClass (for example type: hyperdisk-balanced) for the Gen4 node pool.
    • For existing Persistent Disk volumes, migrate the data to a Hyperdisk volume; the original PD cannot be attached to a Gen4 node.
    • If the workload must keep Persistent Disk, keep it on a PD-capable machine series (for example N2) via node selection. This is a machine-type/disk-type incompatibility, not a capacity problem, so increasing disk size or quota does not help.
  • Hyperdisk Pods become unschedulable when a compute class falls back across VM generations (for example N4 priority, N2 fallback), or one StorageClass must serve mixed generations: a single static disk type in the StorageClass is not compatible with every machine series in the fallback list, so Pods cannot bind their volume on the fallback nodes.

    • Use automated disk type selection: set the StorageClass parameters.type to dynamic with hyperdisk-type, pd-type, and disk-type-preference, plus use-allowed-disk-topology: "true", so GKE selects a compatible disk type per node and schedules Pods only onto nodes that support it. One dynamic StorageClass can then span multiple VM generations (requires the GKE versions noted in the docs).

      Example dynamic StorageClass (GKE 1.35.3-gke.1290000+):

      yaml
      apiVersion: storage.k8s.io/v1
      kind: StorageClass
      metadata:
        name: dynamic-volume
      provisioner: pd.csi.storage.gke.io
      volumeBindingMode: WaitForFirstConsumer
      allowVolumeExpansion: true
      parameters:
        type: dynamic
        pd-type: pd-balanced
        hyperdisk-type: hyperdisk-balanced
        # Preferred storage on nodes that support both PD and Hyperdisk;
        # defaults to hyperdisk-type when omitted.
        disk-type-preference: hyperdisk-type
        # Best practice: schedule Pods only onto nodes that support the disk type.
        use-allowed-disk-topology: "true"
  • Mount stops responding due to the fsGroup setting: A Pod configured with a securityContext.fsGroup on a volume that contains a large number of files makes the kubelet recursively change ownership on every file, which can time out the mount. The symptom is:

    Unable to attach or mount volumes for pod; skipping pod ... timed out waiting for the condition

    Confirm by checking the Pod logs for a Setting volume ownership for ... and fsGroup set entry, then apply one of:

    • Reduce the number of files in the volume.
    • Set securityContext.fsGroupChangePolicy: OnRootMismatch so ownership is only changed when the top-level permissions do not match.
    • Stop using the fsGroup setting if it is not required.
Show full SKILL.md (645 more words)Show less
Disk Performance & Node Storage Pressure
  • Poor disk performance (symptoms such as task dockerd:... blocked for more than 300 seconds, PLEG is not healthy, or slow fs: disk usage scans): the node boot disk is shared across the OS, container images, the overlay filesystem, and disk-backed emptyDir volumes, and performance is shared across all disks of the same type on the node.

    • This commonly affects nodes using standard persistent disks smaller than 200 GB. Increase the disk size or switch to SSD, especially for production.
    • Enable Local SSD for ephemeral storage on node pools whose workloads frequently use emptyDir.
  • Slow disk operations cause Pod creation failures: on affected node versions (GKE 1.18–1.23 before the fixed patch releases), the k8s_node container-runtime logs show failed to reserve container name ... is reserved for ... (containerd issue #4604).

    • Mitigate with restartPolicy: Always or OnFailure in the PodSpec, and increase boot-disk IOPS (larger disk or a faster disk type).
    • The permanent fix is containerd 1.6.0+; upgrade to a GKE version that includes it.
  • StoragePressureDetected (high node storage pressure): node condition StoragePressureRootFileSystem becomes True (for example, Disk /dev/nvme0n1 usage 89% exceeds threshold 85%), caused by excessive emptyDir writes, large image pulls, or accumulating logs.

    • Identify usage with df -h on the affected node (focus on /mnt/stateful_partition and ephemeral mounts).
    • Remediate by using larger boot disks, adding Local SSDs for ephemeral storage, setting appropriate ephemeral-storage requests/limits, and cleaning up unused files/images/logs.
Storage Provisioning & Creation Failures
  • The selected machine type ... has a fixed number of local SSD(s): the Local SSD count specified in EphemeralStorageLocalSsdConfig / LocalNvmeSsdBlockConfig does not match the fixed count included with the machine type.

    • Specify a Local SSD count that matches the machine type. For third-generation machine series, omit the Local SSD count flag and the correct value is configured automatically.
  • Hyperdisk Storage Pools: cluster or node-pool creation fails with ZONE_RESOURCE_POOL_EXHAUSTED (or similar Compute Engine resource errors): the target zone lacks capacity for the requested Hyperdisk Balanced disks or machine type.

    • Select a new zone in the same region that has capacity and where Hyperdisk Balanced Storage Pools are available. Because storage pools are zonal, delete and recreate the pool in the new zone, then create the cluster/node pool there.
Volume Expansion Not Reflecting in the Container

Volume expansion must always be driven through the PersistentVolumeClaim. Editing the PersistentVolume directly can leave the container filesystem on the old size.

  1. Keep the modified PersistentVolume object as it is.

  2. Edit the PersistentVolumeClaim and set spec.resources.requests.storage to a value higher than the current PersistentVolume size.

  3. The kubelet then resizes the PV, PVC, and container filesystem automatically. Verify inside the Pod:

    bash
    kubectl exec {pod_name} -n {workload_namespace} -- df -h
Cloud Storage FUSE Out-Of-Memory (OOM) Events

If Pods experience high memory use or OOM kills related to the Cloud Storage FUSE CSI driver:

  1. Enable CPU/memory snapshots by configuring Cloud Profiler on the Cloud Storage FUSE CSI driver sidecar container.

  2. Locate the OOM event in Cloud Logging, filtering by Pod:

    jsonPayload.involvedObject.name="{pod_name}"
    jsonPayload.involvedObject.kind="Pod"
    OOMKilled

    If the sidecar mounter or GCSFuse process OOMs, the Pod name is the workload Pod's name; if the node driver OOMs, it is gcsfusecsi-node-*.

  3. Extract the Pod UID (jsonPayload.involvedObject.uid) and timestamp, then analyze the matching snapshot in Cloud Profiler using the {pod_name}_{pod_uid} Service Version at the OOM timestamp.


Step 3: Propose the GitOps Correction

Enforce the read-only diagnostics boundary: do not apply live mutations with kubectl edit, kubectl patch, or kubectl apply. Instead, present the corrected StorageClass, PersistentVolumeClaim, PodSpec (securityContext, restartPolicy), or node-pool configuration as a reviewable patch to be applied through the user's GitOps pipeline (for example, Config Sync, Argo CD, or Flux).

References

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/cloud/gke-storage-troubleshooting of google/skills.

Open the folder on GitHubat commit 5120a76

Compare with similar skills

Gke Storage Troubleshooting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gke Storage Troubleshooting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gke Storage Troubleshooting this skillgoogle/skills21k—~3.5kAutomated safety check: PassApache-2.0
Devopsnicepkg/auto-company1942 repos~814Automated safety check: PassMIT
Kcli Cluster Deploymentkarmab/kcli653—~1.5kAutomated safety check: PassApache-2.0
Google Agents CLI Publishpifferologo/cloud-agents-cli1291 repos~2.4kAutomated safety check: PassApache-2.0
DeployingGoogleCloudPlatform/race-condition234—~3kAutomated safety check: PassCustom licence
Aicr Uat ReportNVIDIA/aicr440—~3.2kAutomated safety check: PassApache-2.0

Similar skills

  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    194 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • Guides deployment and management of Kubernetes clusters with kcli.

    653 GitHub stars~1.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Google Agents CLI Publish

    pifferologo/cloud-agents-cli

    This skill should be used when the user wants to "publish an agent", "publish my ADK agent", "register an agent with Gemini Enterprise", "publish to Gemini Enterprise", or needs guidance on the…

    129 GitHub starsUsed in 1 repo~2.4k tokens
    DevOps & CloudAuto-check passed
  • Deploying

    GoogleCloudPlatform/race-condition

    Guides deployment of Race Condition to a GCP project. An agent skill from GoogleCloudPlatform/race-condition.

    234 GitHub stars~3k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • Aicr Uat Report

    NVIDIA/aicr

    Official

    A skill your agent uses when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run…

    440 GitHub stars~3.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • GCP Secret Manager

    sickn33/agentic-awesome-skills

    Secure secrets in Google Cloud Secret Manager. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 3 repos~3.3k tokens
    DevOps & CloudAuto-check passed

More from google/skills

All 147 skills in this repo
  • Official

    Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.

    21k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5k tokensUpdated today
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated today
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated today
    Auto-check passed

Categories

Questions about Gke Storage Troubleshooting

What does Gke Storage Troubleshooting do?

Diagnoses GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs, fsGroup mount timeouts), disk-performance and node storage-pressure issues, slow-disk…. Gke Storage Troubleshooting is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs, fsGroup mount timeouts), disk-performance and node storage-pressure issues, slow-disk Pod-creation failures, volume-expansion problems, Local SSD / Hyperdisk Storage Pool creation errors, and Cloud Storage FUSE OOM.

When should I use Gke Storage Troubleshooting?

Gke Storage Troubleshooting fits situations like: pods are stuck in ContainerCreating; volumes fail to attach; nodes report storage pressure; routine storage provisioning.

How do I install Gke Storage Troubleshooting in Claude Code?

Run `npx skills add google/skills --skill gke-storage-troubleshooting -a claude-code`. Or copy the skill folder (skills/cloud/gke-storage-troubleshooting in google/skills) into .claude/skills/gke-storage-troubleshooting in your project. Claude Code loads it when a task matches its description.

How do I install Gke Storage Troubleshooting in Codex?

Run `npx skills add google/skills --skill gke-storage-troubleshooting -a codex`. Or copy the skill folder (skills/cloud/gke-storage-troubleshooting in google/skills) into .agents/skills/gke-storage-troubleshooting in your project. Codex loads it when a task matches its description.

Can I use Gke Storage Troubleshooting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-storage-troubleshooting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-storage-troubleshooting, .gemini/skills/gke-storage-troubleshooting, .github/skills/gke-storage-troubleshooting and .opencode/skills/gke-storage-troubleshooting in your project.

What does Gke Storage Troubleshooting need to run?

Going by SKILL.md and its folder, Gke Storage Troubleshooting needs the command-line tools its instructions call (kubectl).

Does Gke Storage Troubleshooting access the network?

SKILL.md names 1 domain. As links in the text: docs.cloud.google.com. This is read from the text; nothing was executed.

Is Gke Storage Troubleshooting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gke Storage Troubleshooting use?

Gke Storage Troubleshooting is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gke Storage Troubleshooting use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gke Storage Troubleshooting?

Skills that share tags, products or a category with Gke Storage Troubleshooting: Devops (nicepkg/auto-company, 194 stars), Kcli Cluster Deployment (karmab/kcli, 653 stars), Google Agents CLI Publish (pifferologo/cloud-agents-cli, 129 stars) and Deploying (GoogleCloudPlatform/race-condition, 234 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gke Storage Troubleshooting?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,069 GitHub stars. The repository holds 147 skills in this directory. The repository was last updated on October 9, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.