Official agent skill

Gke AI Troubleshooting Tpu Vbar Oom

by google in google/skills

Diagnoses and prevents vbarcontrolagent segfaults, out-of-memory (OOM) errors, and TPU device initialization failures on TPU v6e nodes in GKE caused by race conditions during TPU device resets or…

OfficialApache-2.0Auto-check passedDevelopment

Install Gke AI Troubleshooting Tpu Vbar Oom

skills CLI
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-vbar-oom -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills gke-ai-troubleshooting-tpu-vbar-oom --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-vbar-oom .claude/skills/gke-ai-troubleshooting-tpu-vbar-oom && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gke-ai-troubleshooting-tpu-vbar-oom
GitHub stars
21k
Token cost
~1.6k tokens
SKILL.md length
530 words
Files
3 (incl. scripts, references)
Skills in repo
150
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnoses and prevents vbarcontrolagent segfaults, out-of-memory (OOM) errors, and TPU device initialization failures on TPU v6e nodes in GKE caused by race conditions during TPU device resets or…

  • Works in 4 steps: Context Acquisition & Time Window… → Check for vbar_control_agent OOMs → Investigate tpu-device-plugin Metrics… → …
  • Troubleshooting vbarcontrolagent crashes
  • SKILL.md covers ⚠️ Prerequisites, 🔍 Diagnostic Workflow, 🛠️ Resolution Workflow and 📋 copypaste checklist
  • Runs Shell scripts from its folder

What it does

Gke AI Troubleshooting Tpu Vbar Oom is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses and prevents vbarcontrolagent segfaults, out-of-memory (OOM) errors, and TPU device initialization failures on TPU v6e nodes in GKE caused by race conditions during TPU device resets or high-frequency metrics polling. Use when troubleshooting vbarcontrolagent crashes, memory cgroup OOMs in serial console logs, tpu-device-plugin metrics checksum corruption errors, or custom TPU metrics collection conflicts on GKE TPU v6e nodes. Don't use for general non-TPU container OOM troubleshooting or standard GKE…

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/failure_signatures.md` and `scripts/validate_queries.sh`).

It sits in Development, covering Async programming. It works with Google Kubernetes Engine and Google Cloud. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • Troubleshooting vbarcontrolagent crashes
  • Memory cgroup OOMs in serial console logs
  • Tpu-device-plugin metrics checksum corruption errors
  • Custom TPU metrics collection conflicts on GKE TPU v6e nodes

Example prompts

  • “Use the gke-ai-troubleshooting-tpu-vbar-oom skill to diagnose and prevents vbarcontrolagent segfaults, out-of-memory (OOM) errors, and TPU device…”
  • “/gke-ai-troubleshooting-tpu-vbar-oom”

Requirements

  • A Bash shell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Context Acquisition & Time Window Definition
  2. Check for vbar_control_agent OOMs
  3. Investigate tpu-device-plugin Metrics Fetch Failures [Low Risk]
  4. Check for Custom Metrics Collection Usage [Low Risk]

What it can do on your machine

Read from SKILL.md and the folder at commit 4b940dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gke AI Troubleshooting Tpu Vbar Oom loads about 1.6k tokens when it runs, and up to ~1.9k if it reads all its reference files. Until then it costs about 146 tokens; SKILL.md has 530 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google/skills at commit 4b940dd, republished under its Apache-2.0 licence (© google). 530 words, ~1,592 tokens.

Download SKILL.mdSave it as .claude/skills/gke-ai-troubleshooting-tpu-vbar-oom/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
gke-ai-troubleshooting-tpu-vbar-oom
description
Diagnoses and prevents vbar_control_agent segfaults, out-of-memory (OOM) errors, and TPU device initialization failures on TPU v6e nodes in GKE caused by race conditions during TPU device resets or high-frequency metrics polling. Use when troubleshooting vbar_control_agent crashes, memory cgroup OOMs in serial console logs, tpu-device-plugin metrics checksum corruption errors, or custom TPU metrics collection conflicts on GKE TPU v6e nodes. Don't use for general non-TPU container OOM troubleshooting or standard GKE node lifecycle operations.
metadata.version
1.0.0
metadata.category
CloudObservabilityAndMonitoring

TPU Connection Failure and VBAR OOM Troubleshooting

Use this skill to systematically diagnose and prevent vbar_control_agent segfaults and Out-Of-Memory (OOM) errors on TPU v6e nodes.

⚠️ Prerequisites

  • Cloud Logging must be enabled for the project.
  • Access to the project and cluster via gcloud or equivalent tool.

🔍 Diagnostic Workflow

Step 0: Context Acquisition & Time Window Definition

Independently gather required context using available GCP/GKE tools or use the provided {variable} placeholders:

  • {project_id}: The GCP Project ID (e.g., customer-ai-project-123).
  • {cluster_name}: The GKE Cluster Name (e.g., tpu-cluster-prod).
  • {node_name}: The Node Name or Instance ID (e.g., tpu-node-1).
  • {workload_name}: The Workload Name / JobSet Name (e.g., my-training-job-456).
  • {namespace}: The Workload Namespace.
  • {issue_time}: The timestamp of the issue (e.g., 2026-04-14T20:00:00Z).
Time Handling & Execution Rules
  1. Window Calculation: If an issue timestamp {issue_time} is provided, calculate the query time window as [{issue_time} - 30m] to [{issue_time} + 30m].
    • Let {start_time} = {issue_time} - 30m
    • Let {end_time} = {issue_time} + 30m
  2. Informational vs. Live Execution: If the user request is informational or query-formulation (e.g. "How can I check...", "How do I determine..."), or if live GCP project resources are not actively targetable, directly output the calculated time window, log names, and Cloud Logging filter templates without attempting live log execution commands.
Step 1: Check for vbar_control_agent OOMs

Look for specific out of memory messages from vbar_control_agent in serial console logs (serialconsole.googleapis.com%2fserial_port_1_output).

  • Tool to use: query_logs (for live diagnostics)
  • Filter Templates:

Serial Console Logs (OOMs):

sql
logName="projects/{project_id}/logs/serialconsole.googleapis.com%2fserial_port_1_output"
AND labels."compute.googleapis.com/resource_name"="{node_name}"
AND SEARCH(text_payload, "Memory cgroup out of memory: Killed process .* (vbar_control_ag)")
AND timestamp >= "{start_time}"
AND timestamp <= "{end_time}"
  • Logic: Presence of Memory cgroup out of memory messages related to vbar_control_agent. Stack traces pointing to libtpu::tpunetd::VBARControlHelper::MetricsReadFromVBAR are a strong indicator.
  • Automation: Proceed to next step automatically after reporting findings.
  • Reference: See references/failure_signatures.md for example log patterns.
Step 2: Investigate tpu-device-plugin Metrics Fetch Failures [Low Risk]

Check if tpu-device-plugin is reporting metric fetch failures.

  • Tool to use: query_logs
  • Filter Template:
sql
resource.type="k8s_container"
AND resource.labels.project_id="{project_id}"
AND resource.labels.cluster_name="{cluster_name}"
AND resource.labels.container_name="tpu-device-plugin"
AND severity=ERROR
AND textPayload:"metrics fetch failed for .* deviceID and .* device path with error: checksum didn't match with the metrics data. Corrupt data found"
AND timestamp >= "{start_time}"
AND timestamp <= "{end_time}"
  • Logic: Errors indicating "metrics fetch failed" with "checksum didn't match" suggest vBAR memory corruption.
  • Automation: Proceed to next step automatically after reporting findings.
Show full SKILL.md (219 more words)Show less
Step 3: Check for Custom Metrics Collection Usage [Low Risk]

Inspect cluster configurations, workloads, or container specs to determine if custom TPU metrics collection mechanisms are deployed.

  • Action: Check if custom scripts or agents (e.g., using libtpu.sdk.tpumonitoring) are deployed that frequently query GetHostMetrics from vBAR Control Agent.

  • Verification Commands:

    • Kubectl Search (Inspect workload env/specs):
    bash
    kubectl get pods -A -o jsonpath='{range .items[*]}{.metadata.namespace}{"/"}{.metadata.name}{"\t"}{.spec.containers[*].image}{"\n"}{end}'
    • Log Search Filter (query_logs):
    sql
    resource.type="k8s_container"
    AND resource.labels.project_id="{project_id}"
    AND resource.labels.cluster_name="{cluster_name}"
    AND textPayload:"libtpu.sdk.tpumonitoring"
    AND timestamp >= "{start_time}"
    AND timestamp <= "{end_time}"
  • Logic: Confirmation of custom metrics collection helps confirm the race condition hypothesis.

🛠️ Resolution Workflow

Resolution 1: Temporarily Disable Custom Metrics Collection [High Risk]

If a custom metrics collection agent is identified, recommend disabling it.

  • Action: Recommend disabling the custom metrics collector.
  • Justification: Prevents reads from vBAR during device resets, stopping crashes and OOMs.
Resolution 2: Await vbar_control_agent Resiliency Update [Low Risk]

Advise that a permanent fix will be available in a future GKE version.

  • Action: Recommend upgrading GKE when the fix is available.
  • Justification: The updated agent will be resilient to memory corruption and gracefully handle reads from unbound vBARs.

📋 copypaste checklist

  • Acquire context and compute [{start_time}, {end_time}] window.
  • Check for vbar_control_agent segfaults and OOMs using query_logs.
  • Investigate tpu-device-plugin failures using query_logs.
  • Inspect for custom metrics collection usage.
  • Advise disabling custom metrics collection if applicable.
  • Advise awaiting resiliency update.

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in skills/cloud/gke-ai-troubleshooting-tpu-vbar-oom of google/skills.

  • SKILL.md
  • references/failure_signatures.md
  • scripts/validate_queries.sh

Open the folder on GitHubat commit 4b940dd

Compare with similar skills

Gke AI Troubleshooting Tpu Vbar Oom next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gke AI Troubleshooting Tpu Vbar Oom compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gke AI Troubleshooting Tpu Vbar Oom this skillgoogle/skills21k—~1.6kAutomated safety check: PassApache-2.0
DeployingGoogleCloudPlatform/race-condition234—~3kAutomated safety check: PassCustom licence
Drawio GCPsparklabx/drawio-ai-kit655—~1.6kAutomated safety check: PassMIT
ContributingGoogleCloudPlatform/race-condition234—~930Automated safety check: PassCustom licence
Exploring The CodebaseGoogleCloudPlatform/race-condition234—~2.6kAutomated safety check: PassCustom licence
Getting StartedGoogleCloudPlatform/race-condition234—~1kAutomated safety check: NotesCustom licence

Similar skills

  • Deploying

    GoogleCloudPlatform/race-condition

    Guides deployment of Race Condition to a GCP project. An agent skill from GoogleCloudPlatform/race-condition.

    234 GitHub stars~3k tokensUpdated 6 days ago
    DevOps & CloudAuto-check passed
  • Drawio GCP

    sparklabx/drawio-ai-kit

    A skill your agent uses when the user asks for a GCP or Google Cloud architecture diagram — VPC/networking, GKE, Cloud Run, landing zone, multi-region, or any diagram built with GCP service icons.

    655 GitHub stars~1.6k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Contributing

    GoogleCloudPlatform/race-condition

    Guides the developer workflow for contributing to Race Condition.

    234 GitHub stars~930 tokensUpdated 6 days ago
    DevelopmentAuto-check passed
  • Exploring The Codebase

    GoogleCloudPlatform/race-condition

    Explains the Race Condition architecture, the design decisions behind it, and where to read code first.

    234 GitHub stars~2.6k tokensUpdated 6 days ago
    DevelopmentAuto-check passed
  • Getting Started

    GoogleCloudPlatform/race-condition

    Guides setup of the Race Condition project from clone to running simulation.

    234 GitHub stars~1k tokensUpdated 6 days ago
    DevelopmentAuto-check: notes
  • Certmanager Dns01 Gke Private Cluster

    divinevideo/divine-mobile

    Fix cert-manager DNS01 ACME challenges stuck in "pending" state with "DNS record not yet propagated" inside GKE private clusters, even when TXT records exist in Cloudflare DNS.

    266 GitHub stars~1.8k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from google/skills

All 150 skills in this repo
  • Official

    Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.

    21k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated yesterday
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Gke AI Troubleshooting Tpu Vbar Oom

What does Gke AI Troubleshooting Tpu Vbar Oom do?

Diagnoses and prevents vbarcontrolagent segfaults, out-of-memory (OOM) errors, and TPU device initialization failures on TPU v6e nodes in GKE caused by race conditions during TPU device resets or…. Gke AI Troubleshooting Tpu Vbar Oom is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses and prevents vbarcontrolagent segfaults, out-of-memory (OOM) errors, and TPU device initialization failures on TPU v6e nodes in GKE caused by race conditions during TPU device resets or high-frequency metrics polling.

When should I use Gke AI Troubleshooting Tpu Vbar Oom?

Gke AI Troubleshooting Tpu Vbar Oom fits situations like: troubleshooting vbarcontrolagent crashes; memory cgroup OOMs in serial console logs; tpu-device-plugin metrics checksum corruption errors; custom TPU metrics collection conflicts on GKE TPU v6e nodes.

How do I install Gke AI Troubleshooting Tpu Vbar Oom in Claude Code?

Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-vbar-oom -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-vbar-oom in google/skills) into .claude/skills/gke-ai-troubleshooting-tpu-vbar-oom in your project. Claude Code loads it when a task matches its description.

How do I install Gke AI Troubleshooting Tpu Vbar Oom in Codex?

Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-vbar-oom -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-vbar-oom in google/skills) into .agents/skills/gke-ai-troubleshooting-tpu-vbar-oom in your project. Codex loads it when a task matches its description.

Can I use Gke AI Troubleshooting Tpu Vbar Oom in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-vbar-oom -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-tpu-vbar-oom, .gemini/skills/gke-ai-troubleshooting-tpu-vbar-oom, .github/skills/gke-ai-troubleshooting-tpu-vbar-oom and .opencode/skills/gke-ai-troubleshooting-tpu-vbar-oom in your project.

What does Gke AI Troubleshooting Tpu Vbar Oom need to run?

Going by SKILL.md and its folder, Gke AI Troubleshooting Tpu Vbar Oom needs a shell for the scripts in its folder. Our summary lists: A Bash shell.

Does Gke AI Troubleshooting Tpu Vbar Oom access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gke AI Troubleshooting Tpu Vbar Oom safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gke AI Troubleshooting Tpu Vbar Oom use?

Gke AI Troubleshooting Tpu Vbar Oom is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gke AI Troubleshooting Tpu Vbar Oom use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 346 tokens, read only when the agent opens those files.

What are the alternatives to Gke AI Troubleshooting Tpu Vbar Oom?

Skills that share tags, products or a category with Gke AI Troubleshooting Tpu Vbar Oom: Deploying (GoogleCloudPlatform/race-condition, 234 stars), Drawio GCP (sparklabx/drawio-ai-kit, 655 stars), Contributing (GoogleCloudPlatform/race-condition, 234 stars) and Exploring The Codebase (GoogleCloudPlatform/race-condition, 234 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gke AI Troubleshooting Tpu Vbar Oom?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.