Official agent skill

Gke AI Troubleshooting Handle Disruption GPU Tpu

by google in google/skills

Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Gke AI Troubleshooting Handle Disruption GPU Tpu

skills CLI
$ npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills gke-ai-troubleshooting-handle-disruption-gpu-tpu --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu .claude/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gke-ai-troubleshooting-handle-disruption-gpu-tpu
GitHub stars
21k
Token cost
~1.6k tokens
SKILL.md length
601 words
Files
1
Skills in repo
150
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE.

  • Works in 5 steps: Context Acquisition → [Low Risk] Check for Upcoming Scheduled… → [Low Risk] Investigation via Cloud… → …
  • Diagnosing node disruptions
  • Calls kubectl
  • Predicting host maintenance events on GPU/TPU nodepools

What it does

Gke AI Troubleshooting Handle Disruption GPU Tpu is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE. Use when diagnosing node disruptions, predicting host maintenance events on GPU/TPU nodepools, inspecting node interruption PromQL metrics, auditing node taints, or configuring workload protection strategies (graceful termination, opportunistic maintenance, PodDisruptionBudgets). Don't use for general GKE cluster creation, network policy…

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud. It works with Google Kubernetes Engine, Prometheus and Kubernetes. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • Diagnosing node disruptions
  • Predicting host maintenance events on GPU/TPU nodepools
  • Inspecting node interruption PromQL metrics
  • Auditing node taints

Example prompts

  • “/gke-ai-troubleshooting-handle-disruption-gpu-tpu”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Context Acquisition
  2. [Low Risk] Check for Upcoming Scheduled Maintenance
  3. [Low Risk] Investigation via Cloud Monitoring (PromQL)
  4. [Low Risk] Investigation via Cloud Logging & Node Taints
  5. Conclusion and Resolution

What it can do on your machine

Read from SKILL.md and the folder at commit 4b940dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.cloud.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gke AI Troubleshooting Handle Disruption GPU Tpu loads about 1.6k tokens when it runs. Until then it costs about 153 tokens; SKILL.md has 601 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~153
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from google/skills at commit 4b940dd, republished under its Apache-2.0 licence (© google). 601 words, ~1,631 tokens.

Download SKILL.mdSave it as .claude/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu/SKILL.md (or your agent's skills folder).
name
gke-ai-troubleshooting-handle-disruption-gpu-tpu
description
Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE. Use when diagnosing node disruptions, predicting host maintenance events on GPU/TPU nodepools, inspecting node interruption PromQL metrics, auditing node taints, or configuring workload protection strategies (graceful termination, opportunistic maintenance, PodDisruptionBudgets). Don't use for general GKE cluster creation, network policy configuration, or non-disruption workload deployment.
metadata.version
1.0.0
metadata.category
CloudObservabilityAndMonitoring

Handle Disruption on GPUs and TPUs Troubleshooting

🔍 Diagnostic Workflow

Step 0: Context Acquisition
  • Mandatory: When a user asks to debug or investigate an actual workload disruption, node crash, or unexpected restart without providing complete cluster details, you MUST immediately halt and request all missing mandatory parameters (project_id, location, cluster_name, timestamp) BEFORE delivering theories or general diagnostic commands. Only skip context acquisition if the user explicitly requests a generic reusable runbook or provides a complete static telemetry/log dump for offline analysis.
  • Optional: node_name, workload_name, workload_namespace, nodepool_name.
Step 1: [Low Risk] Check for Upcoming Scheduled Maintenance
  • Action: Propose running kubectl to check if nodes have the scheduled maintenance label indicating an upcoming disruption.

  • Example Command:

    bash
    kubectl get nodes -l cloud.google.com/scheduled-maintenance-time -L cloud.google.com/scheduled-maintenance-time
  • Interpretation: The SCHEDULED-MAINTENANCE-TIME column shows the Unix epoch time when the VM is scheduled for maintenance. If this label exists, a disruption is guaranteed to occur.

Step 2: [Low Risk] Investigation via Cloud Monitoring (PromQL)
  • Action: Call any available monitoring tool or provide PromQL for manual verification.

  • Mandatory Monitoring Rule: Whenever recommending follow-up monitoring or interruption tracking over time, you MUST explicitly present a PromQL query using the metric kubernetes_io:node_interruption_count filtered by interruption_reason="HW/SW Maintenance". Do not suggest general Cloud Monitoring dashboards or Metrics Explorer without providing this specific PromQL metric expression.

  • Example Query:

    promql
    # Fetch host maintenance events for nodes
    sum by (interruption_type,interruption_reason)( sum_over_time( kubernetes_io:node_interruption_count{monitored_resource="k8s_node", interruption_reason="HW/SW Maintenance"}[${__interval}]))
    promql
    # See the interruption count aggregated by node pool
    sum by (node_pool_name,interruption_type,interruption_reason)( sum_over_time( kubernetes_io:node_pool_interruption_count{monitored_resource="k8s_node_pool", interruption_reason="HW/SW Maintenance", node_pool_name="{nodepool_name}" }[${__interval}]))
  • Interpretation: If kubernetes_io:node_interruption_count shows values > 0 for interruption_reason="HW/SW Maintenance", it indicates the underlying Compute Engine VM was interrupted due to scheduled host maintenance.

Step 3: [Low Risk] Investigation via Cloud Logging & Node Taints
  • Action: Call query_logs or instruct the user to filter their GKE logs for active host maintenance events, and check node taints.
  • Guidance: Look for occurrences in Cloud Logging where cloud.google.com/active-node-maintenance is set to ONGOING. To check if GKE has cordoned the terminating node to prevent new workloads from being scheduled, verify whether the cloud.google.com/impending-node-termination:NoSchedule taint is present (either in GKE event logs or directly via kubectl describe node).
  • Interpretation:
    • cloud.google.com/active-node-maintenance set to ONGOING means workloads are actively being stopped by GKE due to host maintenance.
    • cloud.google.com/impending-node-termination:NoSchedule taint means GKE has cordoned the node to prevent new Pods from being scheduled on the terminating node. DO NOT recommend tolerating this taint.
Show full SKILL.md (212 more words)Show less
Step 4: Conclusion and Resolution
  • Action: Provide a summary of findings to the user and suggest appropriate mitigation strategies if host maintenance events were confirmed or scheduled.
  • Reporting Rule: Signal Only. Report high-signal information indicating that the disruption was caused by Compute Engine host maintenance, specifically affecting the underlying GPU/TPU nodes. DO NOT dump raw logs.
  • Negative Findings Rule-Out: If node scheduled-maintenance labels, PromQL interruption counts, and active maintenance logs all return negative/empty results, definitively conclude that Compute Engine host maintenance did NOT cause the disruption. Direct the user to investigate application-level causes (such as OOMKill events, CUDA runtime errors, or resource limits) and do not propose host maintenance mitigations as the primary resolution.
  • Mandatory Workload Protection Triad: Whenever host maintenance is identified or anticipated on GPU/TPU nodes, consistently recommend all three complementary mitigations together:
    1. Configure Graceful Termination: For workloads that need time to save state (e.g., ML frameworks checkpointing via Orbax), follow the guide to Enable disruption handling and set spec.terminationGracePeriodSeconds (up to 60 minutes) to handle the SIGTERM signal before node shutdown.
    2. Enable Opportunistic Maintenance: To automatically trigger maintenance when GKE detects that GPU/TPU nodes are idle, configure Opportunistic Maintenance.
    3. Configure PodDisruptionBudgets (PDBs): Ensure your workload uses a PodDisruptionBudget to maintain minAvailable replicas during evictions and disruptions.

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu of google/skills.

Open the folder on GitHubat commit 4b940dd

Compare with similar skills

Gke AI Troubleshooting Handle Disruption GPU Tpu next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gke AI Troubleshooting Handle Disruption GPU Tpu compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gke AI Troubleshooting Handle Disruption GPU Tpu this skillgoogle/skills21k—~1.6kAutomated safety check: PassApache-2.0
Mirrord Operatormetalbear-co/mirrord5.4k1 repos~4.6kAutomated safety check: PassMIT
Devopsnicepkg/auto-company1952 repos~814Automated safety check: PassMIT
KubeShark for KubernetesLukasNiessen/kubernetes-skill446—~1.2kAutomated safety check: PassMIT
Kcli Cluster Deploymentkarmab/kcli653—~1.5kAutomated safety check: PassApache-2.0
Aicr Analyzing SnapshotsNVIDIA/aicr440—~3.5kAutomated safety check: PassApache-2.0

Similar skills

  • Mirrord Operator

    metalbear-co/mirrord

    Help users install and configure the mirrord Operator for team/enterprise environments.

    5.4k GitHub starsUsed in 1 repo~4.6k tokens
    DevOps & CloudAuto-check passed
  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    195 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • KubeShark for Kubernetes

    LukasNiessen/kubernetes-skill

    Keeps Kubernetes manifests, Helm charts and policies grounded by diagnosing six failure modes, such as insecure defaults and API drift, and loading only matching references.

    446 GitHub stars~1.2k tokensUpdated 28 days ago
    DevOps & CloudAuto-check passed
  • Guides deployment and management of Kubernetes clusters with kcli.

    653 GitHub stars~1.5k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Official

    A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…

    440 GitHub stars~3.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Signoz

    qjoly/GitOps

    Manage the self-hosted SigNoz observability stack in this GitOps repo.

    112 GitHub stars~6.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from google/skills

All 150 skills in this repo
  • Official

    Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.

    21k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated yesterday
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Gke AI Troubleshooting Handle Disruption GPU Tpu

What does Gke AI Troubleshooting Handle Disruption GPU Tpu do?

Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE. Gke AI Troubleshooting Handle Disruption GPU Tpu is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE.

When should I use Gke AI Troubleshooting Handle Disruption GPU Tpu?

Gke AI Troubleshooting Handle Disruption GPU Tpu fits situations like: diagnosing node disruptions; predicting host maintenance events on GPU/TPU nodepools; inspecting node interruption PromQL metrics; auditing node taints.

How do I install Gke AI Troubleshooting Handle Disruption GPU Tpu in Claude Code?

Run `npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu in google/skills) into .claude/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu in your project. Claude Code loads it when a task matches its description.

How do I install Gke AI Troubleshooting Handle Disruption GPU Tpu in Codex?

Run `npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu in google/skills) into .agents/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu in your project. Codex loads it when a task matches its description.

Can I use Gke AI Troubleshooting Handle Disruption GPU Tpu in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu, .gemini/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu, .github/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu and .opencode/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu in your project.

What does Gke AI Troubleshooting Handle Disruption GPU Tpu need to run?

Going by SKILL.md and its folder, Gke AI Troubleshooting Handle Disruption GPU Tpu needs the command-line tools its instructions call (kubectl).

Does Gke AI Troubleshooting Handle Disruption GPU Tpu access the network?

SKILL.md names 1 domain. As links in the text: docs.cloud.google.com. This is read from the text; nothing was executed.

Is Gke AI Troubleshooting Handle Disruption GPU Tpu safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gke AI Troubleshooting Handle Disruption GPU Tpu use?

Gke AI Troubleshooting Handle Disruption GPU Tpu is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gke AI Troubleshooting Handle Disruption GPU Tpu use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gke AI Troubleshooting Handle Disruption GPU Tpu?

Skills that share tags, products or a category with Gke AI Troubleshooting Handle Disruption GPU Tpu: Mirrord Operator (metalbear-co/mirrord, 5.4k stars), Devops (nicepkg/auto-company, 195 stars), KubeShark for Kubernetes (LukasNiessen/kubernetes-skill, 446 stars) and Kcli Cluster Deployment (karmab/kcli, 653 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gke AI Troubleshooting Handle Disruption GPU Tpu?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.