Devops
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
Diagnoses GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs, fsGroup mount timeouts), disk-performance and node storage-pressure issues, slow-disk…
$ npx skills add google/skills --skill gke-storage-troubleshooting -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills gke-storage-troubleshooting --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-storage-troubleshooting .claude/skills/gke-storage-troubleshooting && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gke-storage-troubleshooting" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-storage-troubleshooting into .claude/skills/gke-storage-troubleshooting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-storage-troubleshooting", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/gke-storage-troubleshootingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill gke-storage-troubleshooting -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills gke-storage-troubleshooting --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/gke-storage-troubleshooting .agents/skills/gke-storage-troubleshooting && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gke-storage-troubleshooting" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-storage-troubleshooting into .agents/skills/gke-storage-troubleshooting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-storage-troubleshooting", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-storage-troubleshooting -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills gke-storage-troubleshooting --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/gke-storage-troubleshooting .cursor/skills/gke-storage-troubleshooting && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gke-storage-troubleshooting" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-storage-troubleshooting into .cursor/skills/gke-storage-troubleshooting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-storage-troubleshooting", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/gke-storage-troubleshooting--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill gke-storage-troubleshooting -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills gke-storage-troubleshooting --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/gke-storage-troubleshooting .gemini/skills/gke-storage-troubleshooting && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gke-storage-troubleshooting" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-storage-troubleshooting into .gemini/skills/gke-storage-troubleshooting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-storage-troubleshooting", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills gke-storage-troubleshootingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill gke-storage-troubleshooting -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/gke-storage-troubleshooting .github/skills/gke-storage-troubleshooting && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gke-storage-troubleshooting" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-storage-troubleshooting into .github/skills/gke-storage-troubleshooting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-storage-troubleshooting", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-storage-troubleshooting -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills gke-storage-troubleshooting --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/gke-storage-troubleshooting .opencode/skills/gke-storage-troubleshooting && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gke-storage-troubleshooting" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-storage-troubleshooting into .opencode/skills/gke-storage-troubleshooting/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-storage-troubleshooting", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gke-storage-troubleshootingDiagnoses GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs, fsGroup mount timeouts), disk-performance and node storage-pressure issues, slow-disk…
Gke Storage Troubleshooting is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs, fsGroup mount timeouts), disk-performance and node storage-pressure issues, slow-disk Pod-creation failures, volume-expansion problems, Local SSD / Hyperdisk Storage Pool creation errors, and Cloud Storage FUSE OOM. Use when Pods are stuck in ContainerCreating, volumes fail to attach or mount, or nodes report storage pressure. Don't use for routine storage provisioning or StorageClass/PVC authoring (see the…
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud. It works with Google Kubernetes Engine and Google Cloud. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 5120a76. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.cloud.google.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gke Storage Troubleshooting loads about 3.5k tokens when it runs. Until then it costs about 140 tokens; SKILL.md has 1,471 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from google/skills at commit 5120a76, republished under its Apache-2.0 licence (© google). 1,471 words, ~3,500 tokens.
.claude/skills/gke-storage-troubleshooting/SKILL.md (or your agent's skills folder).Use this skill to systematically diagnose and resolve persistent-storage failures for workloads running on GKE — volume attach/mount errors, disk performance and node storage pressure, volume expansion, storage-related cluster/node-pool creation errors, and Cloud Storage FUSE memory issues. This skill operates non-interactively and enforces a read-only diagnostics boundary before proposing manifest or configuration corrections.
For routine storage provisioning and StorageClass/PVC authoring, use the
gke-storageskill instead. This skill focuses on failure diagnosis.
Parameter Extraction: Extract required context (project_id,
cluster_name, cluster_location, workload_name, workload_namespace,
pod_name, and the relevant pvc_name / pv_name / node_name)
non-interactively from the user prompt, active SETTINGS.md, or environment
defaults:
workload_namespace to default if omitted.kubectl config current-context or gcloud config get-value project).Cluster Credentials & Fallback Mode:
gcloud container clusters get-credentials {cluster_name} --location {cluster_location} --project {project_id}.kubectl / gcloud diagnostic
commands for the human operator to run.Gather the primary signals, then jump to the matching branch under Step 2 (Resolution) — you normally perform only the one branch that matches your diagnosis, not all of them.
Diagnostic Commands:
kubectl describe pod {pod_name} -n {workload_namespace}
kubectl get pvc,pv -n {workload_namespace}
kubectl get events -n {workload_namespace} --sort-by='.metadata.creationTimestamp'
kubectl describe node {node_name}ContainerCreating with an attach/mount event → Volume
Attach & Mount Failures.PLEG is not healthy, or StoragePressureDetected
events → Disk Performance & Node Storage Pressure.Perform only the branch that matches your Step 1 diagnosis. These branches are mutually exclusive alternatives, not sequential steps.
Error 400: Cannot attach RePD to an optimized VM: Regional persistent
disks are restricted from being used with memory-optimized or
compute-optimized machine types.
Pods stay Pending / FailedScheduling after a node pool is moved to a
4th-generation (N4, N4A, N4D) machine series while the workload uses a
Persistent Disk StorageClass: N4/N4A/N4D machines do not support
Persistent Disk (they support Hyperdisk only), so a PVC bound to a pd-*
StorageClass cannot bind or schedule on those nodes. Events typically show
FailedScheduling with a volume node-affinity / topology conflict.
type: hyperdisk-balanced) for the Gen4 node pool.Hyperdisk Pods become unschedulable when a compute class falls back across VM generations (for example N4 priority, N2 fallback), or one StorageClass must serve mixed generations: a single static disk type in the StorageClass is not compatible with every machine series in the fallback list, so Pods cannot bind their volume on the fallback nodes.
Use automated disk type selection: set the StorageClass
parameters.type to dynamic with hyperdisk-type, pd-type, and
disk-type-preference, plus use-allowed-disk-topology: "true", so GKE
selects a compatible disk type per node and schedules Pods only onto
nodes that support it. One dynamic StorageClass can then span multiple
VM generations (requires the GKE versions noted in the docs).
Example dynamic StorageClass (GKE 1.35.3-gke.1290000+):
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: dynamic-volume
provisioner: pd.csi.storage.gke.io
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
parameters:
type: dynamic
pd-type: pd-balanced
hyperdisk-type: hyperdisk-balanced
# Preferred storage on nodes that support both PD and Hyperdisk;
# defaults to hyperdisk-type when omitted.
disk-type-preference: hyperdisk-type
# Best practice: schedule Pods only onto nodes that support the disk type.
use-allowed-disk-topology: "true"Mount stops responding due to the fsGroup setting: A Pod configured
with a securityContext.fsGroup on a volume that contains a large number
of files makes the kubelet recursively change ownership on every file,
which can time out the mount. The symptom is:
Unable to attach or mount volumes for pod; skipping pod ... timed out waiting for the conditionConfirm by checking the Pod logs for a Setting volume ownership for ... and fsGroup set entry, then apply one of:
securityContext.fsGroupChangePolicy: OnRootMismatch so ownership
is only changed when the top-level permissions do not match.fsGroup setting if it is not required.Poor disk performance (symptoms such as task dockerd:... blocked for more than 300 seconds, PLEG is not healthy, or slow fs: disk usage
scans): the node boot disk is shared across the OS, container images, the
overlay filesystem, and disk-backed emptyDir volumes, and performance is
shared across all disks of the same type on the node.
emptyDir.Slow disk operations cause Pod creation failures: on affected node
versions (GKE 1.18–1.23 before the fixed patch releases), the k8s_node container-runtime logs show failed to reserve container name ... is reserved for ... (containerd issue #4604).
restartPolicy: Always or OnFailure in the PodSpec, and
increase boot-disk IOPS (larger disk or a faster disk type).StoragePressureDetected (high node storage pressure): node condition
StoragePressureRootFileSystem becomes True (for example, Disk /dev/nvme0n1 usage 89% exceeds threshold 85%), caused by excessive
emptyDir writes, large image pulls, or accumulating logs.
df -h on the affected node (focus on
/mnt/stateful_partition and ephemeral mounts).ephemeral-storage requests/limits, and
cleaning up unused files/images/logs.The selected machine type ... has a fixed number of local SSD(s): the
Local SSD count specified in EphemeralStorageLocalSsdConfig /
LocalNvmeSsdBlockConfig does not match the fixed count included with the
machine type.
count flag and
the correct value is configured automatically.Hyperdisk Storage Pools: cluster or node-pool creation fails with
ZONE_RESOURCE_POOL_EXHAUSTED (or similar Compute Engine resource errors):
the target zone lacks capacity for the requested Hyperdisk Balanced disks or
machine type.
Volume expansion must always be driven through the PersistentVolumeClaim. Editing the PersistentVolume directly can leave the container filesystem on the old size.
Keep the modified PersistentVolume object as it is.
Edit the PersistentVolumeClaim and set spec.resources.requests.storage
to a value higher than the current PersistentVolume size.
The kubelet then resizes the PV, PVC, and container filesystem automatically. Verify inside the Pod:
kubectl exec {pod_name} -n {workload_namespace} -- df -hIf Pods experience high memory use or OOM kills related to the Cloud Storage FUSE CSI driver:
Enable CPU/memory snapshots by configuring Cloud Profiler on the Cloud Storage FUSE CSI driver sidecar container.
Locate the OOM event in Cloud Logging, filtering by Pod:
jsonPayload.involvedObject.name="{pod_name}"
jsonPayload.involvedObject.kind="Pod"
OOMKilledIf the sidecar mounter or GCSFuse process OOMs, the Pod name is the workload
Pod's name; if the node driver OOMs, it is gcsfusecsi-node-*.
Extract the Pod UID (jsonPayload.involvedObject.uid) and timestamp, then
analyze the matching snapshot in Cloud Profiler using the
{pod_name}_{pod_uid} Service Version at the OOM timestamp.
Enforce the read-only diagnostics boundary: do not apply live mutations with
kubectl edit, kubectl patch, or kubectl apply. Instead, present the
corrected StorageClass, PersistentVolumeClaim, PodSpec (securityContext,
restartPolicy), or node-pool configuration as a reviewable patch to be applied
through the user's GitOps pipeline (for example, Config Sync, Argo CD, or Flux).
© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/cloud/gke-storage-troubleshooting of google/skills.
Open the folder on GitHubat commit 5120a76
Gke Storage Troubleshooting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gke Storage Troubleshooting this skillgoogle/skills | 21k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| Devopsnicepkg/auto-company | 194 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| Kcli Cluster Deploymentkarmab/kcli | 653 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Google Agents CLI Publishpifferologo/cloud-agents-cli | 129 | 1 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| DeployingGoogleCloudPlatform/race-condition | 234 | — | ~3k | Automated safety check: Pass | Custom licence | |
| Aicr Uat ReportNVIDIA/aicr | 440 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 |
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
pifferologo/cloud-agents-cli
This skill should be used when the user wants to "publish an agent", "publish my ADK agent", "register an agent with Gemini Enterprise", "publish to Gemini Enterprise", or needs guidance on the…
GoogleCloudPlatform/race-condition
Guides deployment of Race Condition to a GCP project. An agent skill from GoogleCloudPlatform/race-condition.
NVIDIA/aicr
A skill your agent uses when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run…
sickn33/agentic-awesome-skills
Secure secrets in Google Cloud Secret Manager. An agent skill from sickn33/agentic-awesome-skills.
google/skills
Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
Works with
Categories
Diagnoses GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs, fsGroup mount timeouts), disk-performance and node storage-pressure issues, slow-disk…. Gke Storage Troubleshooting is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE persistent-storage failures — volume attach/mount errors (Regional PD on optimized VMs, fsGroup mount timeouts), disk-performance and node storage-pressure issues, slow-disk Pod-creation failures, volume-expansion problems, Local SSD / Hyperdisk Storage Pool creation errors, and Cloud Storage FUSE OOM.
Gke Storage Troubleshooting fits situations like: pods are stuck in ContainerCreating; volumes fail to attach; nodes report storage pressure; routine storage provisioning.
Run `npx skills add google/skills --skill gke-storage-troubleshooting -a claude-code`. Or copy the skill folder (skills/cloud/gke-storage-troubleshooting in google/skills) into .claude/skills/gke-storage-troubleshooting in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill gke-storage-troubleshooting -a codex`. Or copy the skill folder (skills/cloud/gke-storage-troubleshooting in google/skills) into .agents/skills/gke-storage-troubleshooting in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-storage-troubleshooting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-storage-troubleshooting, .gemini/skills/gke-storage-troubleshooting, .github/skills/gke-storage-troubleshooting and .opencode/skills/gke-storage-troubleshooting in your project.
Going by SKILL.md and its folder, Gke Storage Troubleshooting needs the command-line tools its instructions call (kubectl).
SKILL.md names 1 domain. As links in the text: docs.cloud.google.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gke Storage Troubleshooting is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Gke Storage Troubleshooting: Devops (nicepkg/auto-company, 194 stars), Kcli Cluster Deployment (karmab/kcli, 653 stars), Google Agents CLI Publish (pifferologo/cloud-agents-cli, 129 stars) and Deploying (GoogleCloudPlatform/race-condition, 234 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,069 GitHub stars. The repository holds 147 skills in this directory. The repository was last updated on October 9, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.