Official agent skill

Gke Reliability

by google in google/skills

Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Gke Reliability

skills CLI
$ npx skills add google/skills --skill gke-reliability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills gke-reliability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-reliability .claude/skills/gke-reliability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gke-reliability
GitHub stars
21k
Token cost
~1.8k tokens
SKILL.md length
419 words
Files
1
Skills in repo
147
Repo updated
First seen
Licence
Apache-2.0

At a glance

Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints.

  • Works in 6 steps: Verify Cluster High Availability → Pod Disruption Budgets (PDBs) → Health Probes → …
  • Configuring GKE workload reliability
  • SKILL.md covers Golden Path Reliability Defaults, Workflows and Best Practices & Production…
  • Calls kubectl and gcloud

What it does

Gke Reliability is an agent skill from google/skills, published by the product's own GitHub organization. Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints. Use when configuring GKE workload reliability, setting up PDBs, or configuring GKE health probes (liveness, readiness, startup). Don't use for generating K8s YAML manifests (use gke-manifest-generation) or disaster recovery and cluster backups (use gke-backup-dr).

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Backup and disaster recovery and Container orchestration. It works with Google Kubernetes Engine and Kubernetes. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • Configuring GKE workload reliability
  • Setting up PDBs
  • Configuring GKE health probes (liveness
  • Generating K8s YAML manifests (use gke-manifest-generation)

Example prompts

  • “Use the gke-reliability skill to improve GKE workload reliability, using PDBs, health probes, and topology spread constraints”
  • “/gke-reliability”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Verify Cluster High Availability
  2. Pod Disruption Budgets (PDBs)
  3. Health Probes
  4. Graceful Shutdown
  5. Topology Spread Constraints
  6. Replicas

What it can do on your machine

Read from SKILL.md and the folder at commit 7d97937. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • gcloud

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl and gcloud, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gke Reliability loads about 1.8k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 419 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from google/skills at commit 7d97937, republished under its Apache-2.0 licence (© google). 419 words, ~1,799 tokens.

Download SKILL.mdSave it as .claude/skills/gke-reliability/SKILL.md (or your agent's skills folder).
name
gke-reliability
description
Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints. Use when configuring GKE workload reliability, setting up PDBs, or configuring GKE health probes (liveness, readiness, startup). Don't use for generating K8s YAML manifests (use gke-manifest-generation) or disaster recovery and cluster backups (use gke-backup-dr).
metadata.version
1.0.1
metadata.category
Containers

GKE Reliability

Routing Note: To generate Kubernetes YAML manifests (Deployment, StatefulSet, Service, ConfigMap, HTTPRoute, PodDisruptionBudget), open gke-manifest-generation/SKILL.md.

This reference covers high availability and reliability configuration for GKE clusters and workloads.

MCP Tools: get_cluster, get_k8s_resource, describe_k8s_resource, apply_k8s_manifest, list_k8s_events

Golden Path Reliability Defaults

SettingGolden Path ValueNotes
Cluster typeRegional (4 zones:Control plane replicated across
: : us-central1-a/b/c/f) : zones :
Upgrade strategySURGE (maxSurge: 1)Rolling upgrades with extra
: : : capacity :
Auto-repairtrueUnhealthy nodes replaced
: : : automatically :
Auto-upgradetrueNodes follow control plane
: : : version :
Release channelREGULARBalanced freshness and stability
Stateful HAEnabledLeader election for stateful
: : : workloads :

Workflows

1. Verify Cluster High Availability
# MCP (preferred)
get_cluster(name="projects/<PROJECT>/locations/<REGION>/clusters/<CLUSTER>",
  readMask="location,locations,nodePools.locations")

# gcloud fallback
gcloud container clusters describe <CLUSTER> --region <REGION> \
  --format="json(location, locations)" \
  --quiet
  • If location is a region (e.g., us-central1), the control plane is regional
  • If locations has multiple entries, nodes span multiple zones
2. Pod Disruption Budgets (PDBs)

PDBs ensure minimum pod availability during voluntary disruptions (node upgrades, autoscaler scale-down).

Check existing PDBs:

# MCP (preferred)
get_k8s_resource(parent="...", resourceType="poddisruptionbudget")

# kubectl fallback
kubectl get pdb --all-namespaces

Create PDB:

yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: my-app-pdb
  namespace: default
spec:
  minAvailable: 2       # Or use maxUnavailable: 1
  selector:
    matchLabels:
      app: my-app

Every production Deployment with 2+ replicas should have a PDB.

3. Health Probes

Every production container should have liveness and readiness probes. Startup probes are recommended for slow-starting apps.

Check existing probes:

# MCP (preferred)
describe_k8s_resource(parent="...", resourceType="deployment", name="<APP>", namespace="<NS>")

# kubectl fallback
kubectl get deployment <APP> -n <NS> -o yaml | grep -E "livenessProbe|readinessProbe|startupProbe"

Recommended probe configuration:

yaml
spec:
  containers:
  - name: app
    livenessProbe:
      httpGet:
        path: /healthz
        port: 8080
      initialDelaySeconds: 15
      periodSeconds: 10
      timeoutSeconds: 2
      failureThreshold: 3
    readinessProbe:
      httpGet:
        path: /readyz
        port: 8080
      initialDelaySeconds: 5
      periodSeconds: 5
      timeoutSeconds: 2
      failureThreshold: 3
    startupProbe:             # For slow-starting apps
      httpGet:
        path: /healthz
        port: 8080
      initialDelaySeconds: 10
      periodSeconds: 5
      timeoutSeconds: 2
      failureThreshold: 30    # 30 * 5s = 150s max startup time
  • Readiness: Determines when a pod can accept traffic
  • Liveness: Determines when to restart a container
  • Startup: Disables liveness/readiness until the app is ready (prevents premature restarts)
4. Graceful Shutdown

Ensure applications handle SIGTERM and drain in-flight requests:

yaml
spec:
  terminationGracePeriodSeconds: 30    # Default; increase for long-running requests
  containers:
  - name: app
    lifecycle:
      preStop:
        exec:
          command: ["/bin/sh", "-c", "sleep 5"]  # Allow LB to deregister
5. Topology Spread Constraints

Distribute pods across zones and nodes to survive failures:

yaml
spec:
  topologySpreadConstraints:
  - maxSkew: 1
    topologyKey: topology.kubernetes.io/zone
    whenUnsatisfiable: DoNotSchedule
    labelSelector:
      matchLabels:
        app: my-app
  - maxSkew: 1
    topologyKey: kubernetes.io/hostname
    whenUnsatisfiable: ScheduleAnyway
    labelSelector:
      matchLabels:
        app: my-app
  • Zone spread (DoNotSchedule): Hard requirement -- pods must be balanced across zones
  • Node spread (ScheduleAnyway): Best-effort -- prefer distribution but don't block scheduling
Show full SKILL.md (169 more words)Show less
6. Replicas
Workload TypeMinimum ReplicasReason
Stateless web/API2Survive single pod/node
: : : failure :
Critical services3Survive zone failure with zone
: : : spread :
Stateful (databases)3 (with replication)Application-level quorum
Batch/jobs1Ephemeral by nature

Best Practices & Production Guidelines

  1. Regional clusters for production: Always use regional clusters to survive zone failures.
  2. PDBs for everything: Every production workload with 2+ replicas needs a PodDisruptionBudget (PDB) to protect against voluntary disruptions.
  3. Probes with Explicit Timeouts: Every production container must have both liveness and readiness probes defined. Always explicitly define initialDelaySeconds, periodSeconds, and timeoutSeconds for all probes. Never rely on the Kubernetes default timeout of 1 second if your application requires more, but always set a strict limit to prevent hanging connections.
  4. Zone spreading: Use topology spread constraints to distribute pods across failure domains (zones and nodes).
  5. Graceful shutdown: Handle SIGTERM and set appropriate terminationGracePeriodSeconds with a preStop sleep hook to allow load balancer deregistration.
  6. Maintenance windows: Schedule upgrades during low-traffic periods (see the gke-upgrades skill).

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/cloud/gke-reliability of google/skills.

Open the folder on GitHubat commit 7d97937

Compare with similar skills

Gke Reliability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gke Reliability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gke Reliability this skillgoogle/skills21k—~1.8kAutomated safety check: PassApache-2.0
Devopsnicepkg/auto-company1922 repos~814Automated safety check: PassMIT
KubeShark for KubernetesLukasNiessen/kubernetes-skill444—~1.2kAutomated safety check: PassMIT
Kcli Cluster Deploymentkarmab/kcli653—~1.5kAutomated safety check: PassApache-2.0
Aicr Analyzing SnapshotsNVIDIA/aicr439—~3.5kAutomated safety check: PassApache-2.0
Kopiur Designhome-operations/kopiur113—~2.4kAutomated safety check: PassAGPL-3.0

Similar skills

  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    192 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • KubeShark for Kubernetes

    LukasNiessen/kubernetes-skill

    Keeps Kubernetes manifests, Helm charts and policies grounded by diagnosing six failure modes, such as insecure defaults and API drift, and loading only matching references.

    444 GitHub stars~1.2k tokensUpdated 25 days ago
    DevOps & CloudAuto-check passed
  • Guides deployment and management of Kubernetes clusters with kcli.

    653 GitHub stars~1.5k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Official

    A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…

    439 GitHub stars~3.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Kopiur Design

    home-operations/kopiur

    Design norms and locked decisions for the Kopiur Kopia-native Kubernetes backup operator (Rust/kube-rs).

    113 GitHub stars~2.4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Supercheck Infrastructure Deployment

    supercheck-io/supercheck

    Work on Supercheck Docker Compose, K3s, Kubernetes manifests, gVisor, OpenTofu/Hetzner, secrets, external services, autoscaling, backups, disaster recovery, DNS/TLS, or production deployment.

    215 GitHub stars~1.4k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes

More from google/skills

All 147 skills in this repo
  • Official

    Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.

    21k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5k tokensUpdated today
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated today
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated today
    Auto-check passed

Categories

Questions about Gke Reliability

What does Gke Reliability do?

Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints. Gke Reliability is an agent skill from google/skills, published by the product's own GitHub organization. Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints.

When should I use Gke Reliability?

Gke Reliability fits situations like: configuring GKE workload reliability; setting up PDBs; configuring GKE health probes (liveness; generating K8s YAML manifests (use gke-manifest-generation).

How do I install Gke Reliability in Claude Code?

Run `npx skills add google/skills --skill gke-reliability -a claude-code`. Or copy the skill folder (skills/cloud/gke-reliability in google/skills) into .claude/skills/gke-reliability in your project. Claude Code loads it when a task matches its description.

How do I install Gke Reliability in Codex?

Run `npx skills add google/skills --skill gke-reliability -a codex`. Or copy the skill folder (skills/cloud/gke-reliability in google/skills) into .agents/skills/gke-reliability in your project. Codex loads it when a task matches its description.

Can I use Gke Reliability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-reliability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-reliability, .gemini/skills/gke-reliability, .github/skills/gke-reliability and .opencode/skills/gke-reliability in your project.

What does Gke Reliability need to run?

Going by SKILL.md and its folder, Gke Reliability needs the command-line tools its instructions call (kubectl and gcloud).

Does Gke Reliability access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gke Reliability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gke Reliability use?

Gke Reliability is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gke Reliability use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gke Reliability?

Skills that share tags, products or a category with Gke Reliability: Devops (nicepkg/auto-company, 192 stars), KubeShark for Kubernetes (LukasNiessen/kubernetes-skill, 444 stars), Kcli Cluster Deployment (karmab/kcli, 653 stars) and Aicr Analyzing Snapshots (NVIDIA/aicr, 439 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gke Reliability?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,032 GitHub stars. The repository holds 147 skills in this directory. The repository was last updated on October 8, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.