Official agent skill

Gke AI Troubleshooting Tpu Metrics Monitoring

by google in google/skills

Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Gke AI Troubleshooting Tpu Metrics Monitoring

skills CLI
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills gke-ai-troubleshooting-tpu-metrics-monitoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring .claude/skills/gke-ai-troubleshooting-tpu-metrics-monitoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gke-ai-troubleshooting-tpu-metrics-monitoring
GitHub stars
21k
Token cost
~1.9k tokens
SKILL.md length
546 words
Files
3 (incl. scripts, references)
Skills in repo
150
Repo updated
First seen
Licence
Apache-2.0

At a glance

Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL.

  • Monitoring TensorCore duty cycle
  • SKILL.md covers Step 0: Mandatory Context and Diagnostic Steps
  • Runs Shell scripts from its folder
  • Multi-host TPU node pool availability

What it does

Gke AI Troubleshooting Tpu Metrics Monitoring is an agent skill from google/skills, published by the product's own GitHub organization. Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating MTTR or MTBI metrics for GKE TPUs. Don't use for general non-TPU GKE workload monitoring or non-metric TPU debugging.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/failure_signatures.md` and `scripts/validate_queries.sh`).

It sits in DevOps & Cloud, covering Monitoring and alerting. It works with Google Kubernetes Engine, Prometheus and Kubernetes. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • Monitoring TensorCore duty cycle
  • Multi-host TPU node pool availability
  • Host maintenance
  • Preemption interruptions

Example prompts

  • “Use the gke-ai-troubleshooting-tpu-metrics-monitoring skill to monitor and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system…”
  • “/gke-ai-troubleshooting-tpu-metrics-monitoring”

Requirements

  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit 4b940dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gke AI Troubleshooting Tpu Metrics Monitoring loads about 1.9k tokens when it runs, and up to ~2.2k if it reads all its reference files. Until then it costs about 110 tokens; SKILL.md has 546 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google/skills at commit 4b940dd, republished under its Apache-2.0 licence (© google). 546 words, ~1,860 tokens.

Download SKILL.mdSave it as .claude/skills/gke-ai-troubleshooting-tpu-metrics-monitoring/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
gke-ai-troubleshooting-tpu-metrics-monitoring
description
Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating MTTR or MTBI metrics for GKE TPUs. Don't use for general non-TPU GKE workload monitoring or non-metric TPU debugging.
metadata.version
1.0.0
metadata.category
CloudObservabilityAndMonitoring

GKE TPU Metrics Monitoring Guide

This skill enables the agent to monitor GKE TPU workloads, nodes, and node pools using GKE system metrics. It helps diagnose if workload interruptions or performance issues are caused by underlying infrastructure.

Step 0: Mandatory Context

Independently gather required context (such as cluster details or node pool names) using available GKE and Cloud tools, or use the provided {variable} placeholders:

  • {project_id}: The GCP Project ID.
  • {cluster_name}: The GKE Cluster Name.
  • {location}: The GKE Cluster Location (region or zone).
  • {node_name}: (Optional) The name of the specific GKE node.
  • {node_pool_name}: (Optional) The name of the GKE node pool.

Diagnostic Steps

Step 1: Verify TPU Runtime Metrics Configuration [Low Risk] [Auto]

Before analyzing runtime metrics, verify that the workload is configured to export them. This ensures the cluster and container environment are set up for automated metric scraping and visibility into accelerator health.

  • Action: Verify that the Pod specification and cluster meet the following prerequisites:
    • containerPort: 8431 exposed on the TPU container (required for Prometheus metric scraping).
    • JAX version 0.4.14 or later if using JAX (earlier versions do not export runtime metrics).
    • GKE version is 1.27.4-gke.900 or later (required for TPU runtime metric support).
    • GKE System Metrics are enabled on the cluster (required for Cloud Monitoring ingestion).
Step 2: Monitor TPU Runtime Metrics [Low Risk] [Auto]

If configured correctly, the following metrics are available in Cloud Monitoring (monitored resources k8s_node and k8s_container):

  • Container Metrics:
    • kubernetes.io/container/accelerator/duty_cycle: Percentage of time over the past sampling period (60 seconds) during which the TensorCores were actively processing on a TPU chip.
    • kubernetes.io/container/accelerator/memory_used: Amount of accelerator memory allocated in bytes.
    • kubernetes.io/container/accelerator/memory_total: Total accelerator memory in bytes.
  • Node Metrics:
    • kubernetes.io/node/accelerator/duty_cycle
    • kubernetes.io/node/accelerator/memory_used
    • kubernetes.io/node/accelerator/memory_total
Step 3: Check Node Status Condition [Low Risk] [Auto]

Query the status condition of GKE nodes (GKE version 1.32.1-gke.1357001 or later).

  • PromQL Query (Check if a specific node is Ready):
    promql
    kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", node_name="{node_name}", condition="Ready", status="True"}
  • PromQL Query (List nodes with non-Ready conditions that are True):
    promql
    kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", condition!="Ready", status="True"}
  • PromQL Query (List nodes that are NOT Ready):
    promql
    kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", condition="Ready", status="False"}
  • PromQL Query (Fleet-wide node status):
    promql
    avg by (condition,status)(avg_over_time(kubernetes_io:node_status_condition{monitored_resource="k8s_node"}[5m]))
Show full SKILL.md (258 more words)Show less
Step 4: Check Node Pool Status [Low Risk] [Auto]

Query the status of multi-host TPU node pools.

  • PromQL Query (Verify if a specific node pool is Running):
    promql
    kubernetes_io:node_pool_status{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}", node_pool_name="{node_pool_name}", status="Running"}
  • PromQL Query (Monitor node pools grouped by status):
    promql
    count by (status)(count_over_time(kubernetes_io:node_pool_status{monitored_resource="k8s_node_pool"}[5m]))
    Possible statuses: Provisioning, Running, Error, Reconciling, Stopping.
Step 5: Check Node Pool Availability [Low Risk] [Auto]

Query if all nodes in a multi-host TPU node pool are available.

  • PromQL Query (Check availability over time):
    promql
    avg by (node_pool_name)(avg_over_time(kubernetes_io:node_pool_multi_host_available{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}"}[5m]))
    Value: 1 (True, all nodes available) or 0 (False, some nodes unavailable).
Step 6: Analyze Node Interruptions [Low Risk] [Auto]

Query the count of interruptions for GKE nodes.

  • PromQL Query (Breakdown of interruptions and causes):
    promql
    sum by (interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node"}[5m]))
    Interruption Types: TerminationEvent, MaintenanceEvent, PreemptionEvent. Interruption Reasons: HostError, Eviction, AutoRepair.
  • PromQL Query (Filter for Host Maintenance events):
    promql
    sum by (interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node", interruption_reason="HW/SW Maintenance"}[5m]))
  • PromQL Query (Interruption count aggregated by node pool):
    promql
    sum by (node_pool_name,interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_pool_interruption_count{monitored_resource="k8s_node_pool", interruption_reason="HW/SW Maintenance", node_pool_name="{node_pool_name}"}[5m]))
Step 7: Calculate Recovery and Interruption Metrics [Low Risk] [Auto]

Calculate Mean Time to Recovery (MTTR) and Mean Time Between Interruptions (MTBI) over the last 7 days.

  • PromQL Query (MTTR - Mean Time to Recovery):
    promql
    sum(sum_over_time(kubernetes_io:node_pool_accelerator_times_to_recover_sum{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}"}[7d])) / sum(sum_over_time(kubernetes_io:node_pool_accelerator_times_to_recover_count{monitored_resource="k8s_node_pool",cluster_name="{cluster_name}"}[7d]))
  • PromQL Query (MTBI - Mean Time Between Interruptions):
    promql
    sum(count_over_time(kubernetes_io:node_memory_total_bytes{monitored_resource="k8s_node", node_name=~"gke-tpu.*|gk3-tpu.*", cluster_name="{cluster_name}"}[7d])) / sum(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node", node_name=~"gke-tpu.*|gk3-tpu.*", cluster_name="{cluster_name}"}[7d]))
Step 8: Monitor TPU Host Metrics [Low Risk] [Auto]

For GKE version 1.28.1-gke.1066000 or later, monitor TPU host performance.

  • Container Metrics:
    • kubernetes.io/container/accelerator/tensorcore_utilization: Current percentage of the TensorCore that is utilized.
    • kubernetes.io/container/accelerator/memory_bandwidth_utilization: Current percentage of the accelerator memory bandwidth that is being used.
  • Node Metrics:
    • kubernetes.io/node/accelerator/tensorcore_utilization
    • kubernetes.io/node/accelerator/memory_bandwidth_utilization

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring of google/skills.

  • SKILL.md
  • references/failure_signatures.md
  • scripts/validate_queries.sh

Open the folder on GitHubat commit 4b940dd

Compare with similar skills

Gke AI Troubleshooting Tpu Metrics Monitoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gke AI Troubleshooting Tpu Metrics Monitoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gke AI Troubleshooting Tpu Metrics Monitoring this skillgoogle/skills21k—~1.9kAutomated safety check: PassApache-2.0
Signozqjoly/GitOps112—~6.1kAutomated safety check: PassWTFPL
Proto Backend Moduleaide-family/moon253—~4.1kAutomated safety check: PassNone
Prometheus GrafanaBagelHole/DevOps-Security-Agent-Skills1.2k—~2.5kAutomated safety check: PassMIT
Grafana Dashboardpando85/kaniop132—~987Automated safety check: PassAGPL-3.0
Alloygrafana/skills282—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Signoz

    qjoly/GitOps

    Manage the self-hosted SigNoz observability stack in this GitOps repo.

    112 GitHub stars~6.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Proto Backend Module

    aide-family/moon

    Implements backend modules from proto definitions for goddess, marksman, and rabbit apps.

    253 GitHub stars~4.1k tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed
  • Prometheus Grafana

    BagelHole/DevOps-Security-Agent-Skills

    Set up metrics collection and visualization with Prometheus and Grafana.

    1.2k GitHub stars~2.5k tokensUpdated 4 mo ago
    DevOps & CloudAuto-check passed
  • Grafana Dashboard

    pando85/kaniop

    Improve and validate the Kaniop Grafana dashboard against repository metrics and the grigri live cluster.

    132 GitHub stars~987 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Alloy

    grafana/skills

    Official

    Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /…

    282 GitHub stars~1.3k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Official

    Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization.

    254 GitHub starsUsed in 2 repos~874 tokens
    DevOps & CloudAuto-check passed

More from google/skills

All 150 skills in this repo
  • Official

    Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.

    21k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated yesterday
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Gke AI Troubleshooting Tpu Metrics Monitoring

What does Gke AI Troubleshooting Tpu Metrics Monitoring do?

Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Gke AI Troubleshooting Tpu Metrics Monitoring is an agent skill from google/skills, published by the product's own GitHub organization. Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL.

When should I use Gke AI Troubleshooting Tpu Metrics Monitoring?

Gke AI Troubleshooting Tpu Metrics Monitoring fits situations like: monitoring TensorCore duty cycle; multi-host TPU node pool availability; host maintenance; preemption interruptions.

How do I install Gke AI Troubleshooting Tpu Metrics Monitoring in Claude Code?

Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring in google/skills) into .claude/skills/gke-ai-troubleshooting-tpu-metrics-monitoring in your project. Claude Code loads it when a task matches its description.

How do I install Gke AI Troubleshooting Tpu Metrics Monitoring in Codex?

Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring in google/skills) into .agents/skills/gke-ai-troubleshooting-tpu-metrics-monitoring in your project. Codex loads it when a task matches its description.

Can I use Gke AI Troubleshooting Tpu Metrics Monitoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-tpu-metrics-monitoring, .gemini/skills/gke-ai-troubleshooting-tpu-metrics-monitoring, .github/skills/gke-ai-troubleshooting-tpu-metrics-monitoring and .opencode/skills/gke-ai-troubleshooting-tpu-metrics-monitoring in your project.

What does Gke AI Troubleshooting Tpu Metrics Monitoring need to run?

Going by SKILL.md and its folder, Gke AI Troubleshooting Tpu Metrics Monitoring needs a shell for the scripts in its folder. Our summary lists: A Bash shell.

Does Gke AI Troubleshooting Tpu Metrics Monitoring access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gke AI Troubleshooting Tpu Metrics Monitoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gke AI Troubleshooting Tpu Metrics Monitoring use?

Gke AI Troubleshooting Tpu Metrics Monitoring is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gke AI Troubleshooting Tpu Metrics Monitoring use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 361 tokens, read only when the agent opens those files.

What are the alternatives to Gke AI Troubleshooting Tpu Metrics Monitoring?

Skills that share tags, products or a category with Gke AI Troubleshooting Tpu Metrics Monitoring: Signoz (qjoly/GitOps, 112 stars), Proto Backend Module (aide-family/moon, 253 stars), Prometheus Grafana (BagelHole/DevOps-Security-Agent-Skills, 1.2k stars) and Grafana Dashboard (pando85/kaniop, 132 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gke AI Troubleshooting Tpu Metrics Monitoring?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.