Official agent skill

Gke AI Troubleshooting Tpu Dynamic Slices Monitoring

by google in google/skills

Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Gke AI Troubleshooting Tpu Dynamic Slices Monitoring

skills CLI
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills gke-ai-troubleshooting-tpu-dynamic-slices-monitoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring .claude/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gke-ai-troubleshooting-tpu-dynamic-slices-monitoring
GitHub stars
21k
Token cost
~2k tokens
SKILL.md length
794 words
Files
3 (incl. scripts, references)
Skills in repo
150
Repo updated
First seen
Licence
Apache-2.0

At a glance

Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources.

  • Works in 3 steps: Context Acquisition & Time Window… → Describe the Slice Custom Resource [Low… → Verify Workload Specification [Low Risk]
  • Checking TPU slice lifecycle states
  • SKILL.md covers Prerequisites, Diagnostic Workflow and Resolution & Management Workflow
  • Runs Shell scripts from its folder; calls kubectl

What it does

Gke AI Troubleshooting Tpu Dynamic Slices Monitoring is an agent skill from google/skills, published by the product's own GitHub organization. Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources. Use when checking TPU slice lifecycle states, troubleshooting slice provisioning failures, validating single-slice or multi-slice (JobSet) workload manifests, or safely patching stuck finalizers and disabling the slice controller. Don't use for generic GKE cluster node pool creation or standard non-TPU workload management (use gke-basics or gke-cluster-creation instead).

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/failure_signatures.md` and `scripts/validate_queries.sh`).

It sits in DevOps & Cloud. It works with Google Kubernetes Engine. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • Checking TPU slice lifecycle states
  • Troubleshooting slice provisioning failures
  • Validating single-slice
  • Multi-slice (JobSet) workload manifests

Example prompts

  • “/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring”

Requirements

  • A Bash shell

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Context Acquisition & Time Window Definition
  2. Describe the Slice Custom Resource [Low Risk]
  3. Verify Workload Specification [Low Risk]

What it can do on your machine

Read from SKILL.md and the folder at commit 4b940dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gke AI Troubleshooting Tpu Dynamic Slices Monitoring loads about 2k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 126 tokens; SKILL.md has 794 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google/skills at commit 4b940dd, republished under its Apache-2.0 licence (© google). 794 words, ~1,958 tokens.

Download SKILL.mdSave it as .claude/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
gke-ai-troubleshooting-tpu-dynamic-slices-monitoring
description
Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources. Use when checking TPU slice lifecycle states, troubleshooting slice provisioning failures, validating single-slice or multi-slice (JobSet) workload manifests, or safely patching stuck finalizers and disabling the slice controller. Don't use for generic GKE cluster node pool creation or standard non-TPU workload management (use gke-basics or gke-cluster-creation instead).
metadata.version
1.0.0
metadata.category
Containers

GKE TPU Dynamic Slices Monitoring & Management

Monitors the status of TPU Slice custom resources, troubleshoots provisioning failures, validates workload manifests on dynamic slices, and performs cleanups.

Prerequisites

  • Cloud Logging enabled for the project.
  • kubectl and gcloud CLIs configured to access the GKE cluster.

Diagnostic Workflow

Step 0: Context Acquisition & Time Window Definition

Gather project, cluster, and slice context using cluster tools or the following parameters:

  • Project ID: {project_id} (e.g., my-gcp-project)
  • Cluster Name: {cluster_name} (e.g., tpu-cluster)
  • Region/Zone: {location} (e.g., us-central1-a)
  • Slice Name: {slice_name} (e.g., test-slice)
  • Issue Time: {timestamp} (Optional; default to the last 30 minutes window [T - 30m] to [T + 30m])

Step 1: Describe the Slice Custom Resource [Low Risk]

When asked to inspect, troubleshoot, or check a slice status, immediately execute kubectl describe slice {slice_name} using available cluster tools to perform the inspection. Parse the resulting Status.Conditions output against the condition table below to diagnose the exact state and provide concrete recommendations.

  • Command:

    bash
    kubectl describe slice {slice_name}
State & Reason Analysis

Analyze the Status.Conditions (especially Type: Ready and its Reason and Status):

Lifecycle State / ReasonMeaningRecommended Action
SliceNotCreatedGKE Slice Controller is initializing the slice and performing resource checks.Wait a few minutes and re-check slice status.
SliceCreationFailedPrerequisites validation failed (e.g., selected nodes don't exist, nodes are already used by another slice, or the topology doesn't match the number of partitions).Verify selected nodes exist, are unallocated, and topology matches partition count.
ACTIVATINGGKE is actively forming and provisioning the TPU slice.Monitor node provisioning.
ACTIVEThe TPU slice is successfully formed and ready to host workloads.Proceed to deploy or check workloads.
ACTIVE_DEGRADEDThe slice is usable, but one or more sub-blocks are degraded.Monitor workload logs for interconnect or device errors. Check faulty node VMs.
FAILEDGKE failed to form the TPU slice (e.g., selected nodes are not part of the same reservation block).Ensure all selected nodes belong to the same reservation block.
DEACTIVATINGThe slice is dismantling (triggered by user deletion or a critical systemic failure).Wait for dismantling to finish, or patch finalizers if stuck.
INCOMPLETEThe terminal phase before the Slice CR is deleted from the cluster.No action required; the resource will be removed shortly.
Provisioning Failure Troubleshooting Checklist

When investigating slice creation or provisioning failures (SliceCreationFailed or FAILED), perform the following verification steps:

  1. Node Existence & Allocation Check: Verify that the selected TPU nodes exist in the cluster and are not already allocated to another slice (kubectl get nodes -l cloud.google.com/gke-tpu-slice, kubectl get slice -A).
  2. Topology Alignment: Confirm that the partition count matches the requested topology dimensions (e.g. topology 2x2 requires 4 nodes).
  3. Reservation Block Alignment Check: Confirm that all selected TPU nodes belong to the same reservation and reservation block.

Step 2: Verify Workload Specification [Low Risk]

Ensure workload manifests are configured correctly to target the dynamic slice.

Show full SKILL.md (323 more words)Show less
1. Single-Slice Workload Requirements

Check that the Pod template contains the following annotations and selectors:

  • Annotations:
    • cloud.google.com/gke-tpu-slice-topology: "{topology}" (e.g., "4x4x4")
  • NodeSelector:
    • cloud.google.com/gke-tpu-topology: "{topology}" (e.g., "4x4x4")
    • cloud.google.com/gke-tpu-accelerator: "{accelerator_type}" (e.g., "tpu7x")
    • cloud.google.com/gke-tpu-slice: "{slice_name}" (e.g., "test-slice")
2. Multi-Slice (JobSet) Workload Requirements

If deploying a multi-slice JobSet, verify:

  • JobSet Annotation:
    • alpha.jobset.sigs.k8s.io/exclusive-topology: cloud.google.com/gke-tpu-slice
  • Pod Template Annotations:
    • cloud.google.com/gke-tpu-slice-topology: "{topology}"
  • Pod Template NodeSelector:
    • cloud.google.com/gke-tpu-topology: "{topology}"
    • cloud.google.com/gke-tpu-accelerator: "{accelerator_type}"
    • Note: Do NOT manually specify cloud.google.com/gke-tpu-slice in the nodeSelector; JobSet handles slice assignment automatically.

Resolution & Management Workflow

Resolution 1: Force Delete a Stuck Slice [High Risk]

If a slice is stuck in DEACTIVATING or deletion hangs indefinitely due to stuck finalizers:

  1. Identify Cause: Explain that finalizers on the slice resource (metadata.finalizers) are preventing Kubernetes from completing resource deletion.

  2. Propose Resolution: Propose removing finalizers from the metadata path (/metadata/finalizers) using a JSON patch operation:

    bash
    kubectl patch slice {slice_name} --type json -p='[{"op": "remove", "path": "/metadata/finalizers"}]'
  3. Provide Warning: Explicitly warn the user that removing finalizers bypasses standard controller dismantling and may leave underlying VM, network, or accelerator resources uncleaned or orphaned.

  4. CRITICAL SAFETY MANDATE: The response MUST explicitly ask the user for confirmation (e.g. "Removing finalizers on /metadata/finalizers via JSON patch is a high-risk operation that may leave orphaned resources. Do you confirm you want to apply this patch to slice {slice_name}?") and pause for user confirmation before applying or executing the patch.


Resolution 2: Disable and Clean Up Slice Controller [High Risk]

If dynamic slicing needs to be disabled:

  1. Check for existing Slices:

    bash
    kubectl get slice -A

    Ensure all slices are deleted before disabling the controller.

  2. Disable Slice Controller via gcloud:

    bash
    gcloud container clusters update {cluster_name} \
        --location={location} \
        --no-enable-slice-controller
  3. Delete the Slice CRD:

    bash
    kubectl delete crd slices.accelerator.gke.io
  4. Clean up Node Labels: Remove GKE TPU Slice labels from all nodes in the cluster:

    bash
    kubectl label nodes --all cloud.google.com/gke-tpu-slice- cloud.google.com/gke-tpu-slice-topology-
  • Safety Rule: Propose the exact commands and confirm before executing disabling or destructive cleanup steps.

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring of google/skills.

  • SKILL.md
  • references/failure_signatures.md
  • scripts/validate_queries.sh

Open the folder on GitHubat commit 4b940dd

Compare with similar skills

Gke AI Troubleshooting Tpu Dynamic Slices Monitoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gke AI Troubleshooting Tpu Dynamic Slices Monitoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gke AI Troubleshooting Tpu Dynamic Slices Monitoring this skillgoogle/skills21k—~2kAutomated safety check: PassApache-2.0
Mirrord Operatormetalbear-co/mirrord5.4k1 repos~4.6kAutomated safety check: PassMIT
Devopsnicepkg/auto-company1952 repos~814Automated safety check: PassMIT
KubeShark for KubernetesLukasNiessen/kubernetes-skill446—~1.2kAutomated safety check: PassMIT
Kcli Cluster Deploymentkarmab/kcli653—~1.5kAutomated safety check: PassApache-2.0
Aicr Analyzing SnapshotsNVIDIA/aicr440—~3.5kAutomated safety check: PassApache-2.0

Similar skills

  • Mirrord Operator

    metalbear-co/mirrord

    Help users install and configure the mirrord Operator for team/enterprise environments.

    5.4k GitHub starsUsed in 1 repo~4.6k tokens
    DevOps & CloudAuto-check passed
  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    195 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • KubeShark for Kubernetes

    LukasNiessen/kubernetes-skill

    Keeps Kubernetes manifests, Helm charts and policies grounded by diagnosing six failure modes, such as insecure defaults and API drift, and loading only matching references.

    446 GitHub stars~1.2k tokensUpdated 28 days ago
    DevOps & CloudAuto-check passed
  • Guides deployment and management of Kubernetes clusters with kcli.

    653 GitHub stars~1.5k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Official

    A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…

    440 GitHub stars~3.5k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Google Agents CLI Deploy

    pifferologo/cloud-agents-cli

    This skill should be used when the user wants to "deploy an agent", "deploy my ADK agent", "set up CI/CD", "configure secrets", "troubleshoot a deployment", or needs guidance on Agent Runtime, Cloud…

    129 GitHub starsUsed in 1 repo~6k tokens
    DevOps & CloudAuto-check passed

More from google/skills

All 150 skills in this repo
  • Official

    Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.

    21k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated yesterday
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Gke AI Troubleshooting Tpu Dynamic Slices Monitoring

What does Gke AI Troubleshooting Tpu Dynamic Slices Monitoring do?

Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources. Gke AI Troubleshooting Tpu Dynamic Slices Monitoring is an agent skill from google/skills, published by the product's own GitHub organization. Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources.

When should I use Gke AI Troubleshooting Tpu Dynamic Slices Monitoring?

Gke AI Troubleshooting Tpu Dynamic Slices Monitoring fits situations like: checking TPU slice lifecycle states; troubleshooting slice provisioning failures; validating single-slice; multi-slice (JobSet) workload manifests.

How do I install Gke AI Troubleshooting Tpu Dynamic Slices Monitoring in Claude Code?

Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring in google/skills) into .claude/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring in your project. Claude Code loads it when a task matches its description.

How do I install Gke AI Troubleshooting Tpu Dynamic Slices Monitoring in Codex?

Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring in google/skills) into .agents/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring in your project. Codex loads it when a task matches its description.

Can I use Gke AI Troubleshooting Tpu Dynamic Slices Monitoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring, .gemini/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring, .github/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring and .opencode/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring in your project.

What does Gke AI Troubleshooting Tpu Dynamic Slices Monitoring need to run?

Going by SKILL.md and its folder, Gke AI Troubleshooting Tpu Dynamic Slices Monitoring needs a shell for the scripts in its folder and the command-line tools its instructions call (kubectl). Our summary lists: A Bash shell.

Does Gke AI Troubleshooting Tpu Dynamic Slices Monitoring access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gke AI Troubleshooting Tpu Dynamic Slices Monitoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gke AI Troubleshooting Tpu Dynamic Slices Monitoring use?

Gke AI Troubleshooting Tpu Dynamic Slices Monitoring is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gke AI Troubleshooting Tpu Dynamic Slices Monitoring use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 406 tokens, read only when the agent opens those files.

What are the alternatives to Gke AI Troubleshooting Tpu Dynamic Slices Monitoring?

Skills that share tags, products or a category with Gke AI Troubleshooting Tpu Dynamic Slices Monitoring: Mirrord Operator (metalbear-co/mirrord, 5.4k stars), Devops (nicepkg/auto-company, 195 stars), KubeShark for Kubernetes (LukasNiessen/kubernetes-skill, 446 stars) and Kcli Cluster Deployment (karmab/kcli, 653 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gke AI Troubleshooting Tpu Dynamic Slices Monitoring?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.