Official agent skill

Gke AI Troubleshooting Jobset Interruption

by google in google/skills

Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Gke AI Troubleshooting Jobset Interruption

skills CLI
$ npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills gke-ai-troubleshooting-jobset-interruption --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-jobset-interruption .claude/skills/gke-ai-troubleshooting-jobset-interruption && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gke-ai-troubleshooting-jobset-interruption
GitHub stars
21k
Token cost
~2.8k tokens
SKILL.md length
1,003 words
Files
3 (incl. scripts, references)
Skills in repo
150
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously.

  • Works in 5 steps: Context Acquisition & Time Window… → Identify JobSet Restarts and Attempts… → Inspect Nodepool Interruptions [Low Risk] → …
  • Troubleshooting JobSet restart loops
  • SKILL.md covers ⚠️ Prerequisites & Sandbox Rules, 🔍 Diagnostic Workflow, 🛠️ Resolution Workflow and 📋 Copypaste Checklist
  • Runs Shell scripts from its folder

What it does

Gke AI Troubleshooting Jobset Interruption is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously. Use when troubleshooting JobSet restart loops, spot VM preemptions, node readiness failures, host VM issues, or coordinator worker crashes. Don't use for general GKE cluster creation, basic workload deployment, or non-JobSet application issues.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/failure_signatures.md` and `scripts/validate_queries.sh`).

It sits in DevOps & Cloud. It works with Google Kubernetes Engine and Prometheus. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • Troubleshooting JobSet restart loops
  • Spot VM preemptions
  • Node readiness failures
  • Coordinator worker crashes

Example prompts

  • “Use the gke-ai-troubleshooting-jobset-interruption skill to diagnose GKE JobSet interruptions, restarts, and preemptions for AI/ML training…”
  • “/gke-ai-troubleshooting-jobset-interruption”

Requirements

  • A Bash shell

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Context Acquisition & Time Window Definition
  2. Identify JobSet Restarts and Attempts [Low Risk]
  3. Inspect Nodepool Interruptions [Low Risk]
  4. Inspect Nodes and Underlying Host VMs [Low Risk]
  5. Inspect Pod and Worker / Container Failures [Low Risk]

What it can do on your machine

Read from SKILL.md and the folder at commit 4b940dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gke AI Troubleshooting Jobset Interruption loads about 2.8k tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 98 tokens; SKILL.md has 1,003 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google/skills at commit 4b940dd, republished under its Apache-2.0 licence (© google). 1,003 words, ~2,759 tokens.

Download SKILL.mdSave it as .claude/skills/gke-ai-troubleshooting-jobset-interruption/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
gke-ai-troubleshooting-jobset-interruption
description
Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously. Use when troubleshooting JobSet restart loops, spot VM preemptions, node readiness failures, host VM issues, or coordinator worker crashes. Don't use for general GKE cluster creation, basic workload deployment, or non-JobSet application issues.
metadata.version
1.0.0
metadata.category
Containers

GKE JobSet Interruption Troubleshooting

Use this skill to systematically diagnose and resolve JobSet interruptions, restarts, and preemptions on GKE clusters hosting large-scale AI/ML workloads.

⚠️ Prerequisites & Sandbox Rules

  • JobSet metrics package must be enabled in kube-state-metrics for your cluster.
  • Cloud Logging and Cloud Monitoring enabled for the Google Cloud Project.
  • Autonomous & Sandbox Execution Rule: If API queries, tools, or commands encounter 403 Permission Denied, authentication errors, or network isolation, do NOT enter authentication or credential troubleshooting loops. Populate the query templates with the acquired variables ({project_id}, {cluster_name}, {workload_name}, {start_time}, {end_time}), inspect any locally staged telemetry or mock data files if available, and complete the diagnostic workflow and resolution recommendations autonomously.

🔍 Diagnostic Workflow

Step 0: Context Acquisition & Time Window Definition

Independently gather context using tools, workspace files, environment details, or user prompt context:

  • Project ID ({project_id})
  • Cluster Name ({cluster_name})
  • Workload Name (JobSet Name) ({workload_name})
  • Workload Namespace ({namespace})
  • Issue Time ({issue_time})

If specific variables are not explicitly provided by the user, inspect cluster resources or logs to determine them, or use the {variable} placeholders provided.

Time Handling Rules
  1. Autonomous Time Window: If a relative time (e.g., "X minutes ago") or no exact timestamp is provided, calculate the query window based on current time or available log timestamps.
  2. Window Calculation: If a timestamp {issue_time} is available (or calculated as T), set {start_time} = T - 30m and {end_time} = T + 30m.

Step 1: Identify JobSet Restarts and Attempts [Low Risk]

Verify if the JobSet is experiencing restart loops and determine the frequency of restarts.

Visual Chart / MQL Query - restarts
  • MQL Query Specification:

    mql
    fetch prometheus_target
    | metric 'prometheus.googleapis.com/kube_jobset_restarts/gauge'
    | filter resource.cluster_name == '{cluster_name}' && metric.jobset_name == '{workload_name}'
    | align next_older(1m)
    | every 1m
    | group_by [metric.jobset_name], [val: max(value)]
PromQL Metric Query - restarts
  • PromQL Query Specification:

    promql
    kube_jobset_restarts{jobset_name="{workload_name}", cluster="{cluster_name}"}
  • Diagnostic Logic: A non-zero or increasing value for restarts indicates that the JobSet is being actively restarted by the controller due to worker failure or interruption.

  • Automation: Proceed to Step 2 automatically after reporting findings.


Step 2: Inspect Nodepool Interruptions [Low Risk]

Determine if the JobSet restarts were triggered by physical nodepool-level events (such as spot preemptions, maintenance, or host terminations).

A. Metrics Query (Nodepool Interruption Counts)
Visual Chart / MQL Query - interruptions
  • MQL Query Specification:

    mql
    fetch k8s_node_pool
    | metric 'kubernetes.io/node_pool/interruption_count'
    | filter cluster_name == '{cluster_name}'
    | align next_older(10m)
    | every 10m
    | group_by [metric.interruption_type, metric.interruption_reason, metadata.system.node_pool_name], [val: sum(value)]
PromQL Query - interruptions
  • PromQL Query Specification:

    promql
    sum by (interruption_type, interruption_reason, node_pool_name, cluster_name) (
      avg_over_time(kubernetes_io:node_pool_interruption_count{cluster_name="{cluster_name}"}[10m])
    )
B. Log Query (Nodepool Life Events)
  • LQL Log Filter Specification:

    sql
    resource.type="gke_nodepool"
    AND resource.labels.cluster_name="{cluster_name}"
    AND timestamp >= "{start_time}"
    AND timestamp <= "{end_time}"
  • Diagnostic Logic:

    • PreemptionEvent: Spot VMs were preempted, or node was scale-down.
    • MaintenanceEvent: Node pool updated or Google scheduled maintenance.
    • TerminationEvent: Serious host failures. Check interruption_reason or logs for host issues.
    • See Failure Signatures for examples of node termination logs and preemption events.
  • Automation: Proceed to Step 3 automatically.


Step 3: Inspect Nodes and Underlying Host VMs [Low Risk]

Correlate node readiness failures with physical host VMs to see if a single faulty host repeatedly fails coordinator pods.

A. Metrics Query (Node Ready Status Check)
Visual Chart / MQL Query - node status
  • MQL Query Specification:

    mql
    fetch k8s_node
    | metric 'kubernetes.io/node/status_condition'
    | filter cluster_name == '{cluster_name}' && metric.condition == 'Ready' && metric.status == 'False'
    | align next_older(1m)
    | every 1m
    | group_by [node_name, metadata.user.gke_nodepool], [val: max(value)]
PromQL Query - node status
  • PromQL Query Specification:

    promql
    sum by (status, condition, node_pool_name) (
      kubernetes_io:node_status_condition{cluster_name="{cluster_name}", condition="Ready", status="False"}
    )
B. Metrics Query (Node-to-Host Metadata Topology Correlation)
  • MQL Query Specification:

    mql
    fetch k8s_node
    | metric 'kubernetes.io/node/cpu/total_cores'
    | filter cluster_name == '{cluster_name}'
    | align next_older(1m)
    | every 1m
    | group_by [node_name, metadata.user.gce_topology_host, metadata.user.gke_nodepool], [val: max(value)]
Show full SKILL.md (433 more words)Show less
C. Log Query (Node Fault Logs)
  • LQL Log Filter Specification:

    sql
    resource.type="k8s_node"
    AND resource.labels.cluster_name="{cluster_name}"
    AND (textPayload:"host error" OR textPayload:"kernel panic" OR textPayload:"hardware failure" OR textPayload:"NodeNotReady")
    AND timestamp >= "{start_time}"
    AND timestamp <= "{end_time}"
  • Diagnostic Logic: Identify if specific nodes are unhealthy (Ready=False or Unknown) and correlate them to their GCE physical host ID via metadata.user.gce_topology_host. Check if the same host is repeatedly failing.

  • Automation: Proceed to Step 4 automatically.


Step 4: Inspect Pod and Worker / Container Failures [Low Risk]

Analyze pod status phases and retrieve coordinator worker logs to identify application-level crashes or network deadlocks.

Required Execution Order: You MUST analyze pod status phases (Section A) and unschedulable pod metrics (Section B) to assess overall workload health before inspecting specific worker container logs (Section C).

A. Metrics Query (Pod Lifecycle Phases)
Visual Chart / MQL Query - pod phase
  • MQL Query Specification:

    mql
    fetch k8s_pod
    | metric 'kubernetes.io/pod/status/phase'
    | filter cluster_name == '{cluster_name}' && pod_name ==~ '{workload_name}.*'
    | align next_older(10m)
    | every 10m
    | group_by [metric.phase], [val: count()]
PromQL Query - pod phase
  • PromQL Query Specification:

    promql
    sum by (phase) (
      avg_over_time(kube_pod_status_phase{cluster="{cluster_name}", pod=~"{workload_name}.*"}[10m])
    )
B. Metrics Query (Unschedulable Pod Count)
  • MQL Query Specification:

    mql
    fetch k8s_pod
    | metric 'kubernetes.io/pod/status/unschedulable'
    | filter cluster_name == '{cluster_name}' && pod_name ==~ '{workload_name}.*'
    | align next_older(10m)
    | every 10m
    | group_by [pod_name], [val: max(value)]
C. Log Query (Worker Container Logs)
  • LQL Log Filter Specification:

    sql
    resource.type="k8s_container"
    AND resource.labels.cluster_name="{cluster_name}"
    AND labels."k8s-pod/jobset_sigs_k8s_io/jobset-name"="{workload_name}"
    AND timestamp >= "{start_time}"
    AND timestamp <= "{end_time}"
  • Diagnostic Logic:

    1. Check the pod timeline to spot pending or unschedulable pods.
    2. Use worker container logs to analyze worker 0 in slice 0 (coordinator) for NCCL timeouts, collective communication issues, or MegaScale hangs.
  • Automation: Proceed to Resolution.


🛠️ Resolution Workflow

Resolution 1: Preemption & Autoscaling Optimizations [Low Risk]

If Step 2 showed high preemption counts on Spot VMs:

  • Action: Suggest switching critical long-running training workloads to GKE Reserved/On-Demand VMs or utilizing Compact Placement Policies to minimize defragmentation interruptions.
  • Justification: Eliminates spot-market preemptions and reduces training restarts.
Resolution 2: Quarantine Faulty Host VMs [High Risk]

If Step 3 identified a specific host ID (gce-topology-host) that consistently fails or triggers restarts across multiple attempts:

  • Action: Recommend cordoning/draining the GKE node, deleting the underlying GCE VM instance to trigger instance recreation, and opening a support ticket with Google Cloud Support specifying the physical host ID.
  • Justification: GKE auto-repair will recreate the VM instance on healthy physical hardware, preventing infinite restart loops.

📋 Copypaste Checklist

  • Gather context and compute {start_time} ({issue_time} - 30m) and {end_time} ({issue_time} + 30m) window.
  • Query JobSet restart attempts.
  • Check Nodepool interruptions (spot preemptions vs. hardware terminations).
  • Query node-to-host mapping and check node logs for physical host errors.
  • Inspect pod timeline status and coordinator worker container logs.
  • Recommend appropriate scheduling strategy (On-demand vs Spot) or host VM quarantining.

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in skills/cloud/gke-ai-troubleshooting-jobset-interruption of google/skills.

  • SKILL.md
  • references/failure_signatures.md
  • scripts/validate_queries.sh

Open the folder on GitHubat commit 4b940dd

Compare with similar skills

Gke AI Troubleshooting Jobset Interruption next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gke AI Troubleshooting Jobset Interruption compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gke AI Troubleshooting Jobset Interruption this skillgoogle/skills21k—~2.8kAutomated safety check: PassApache-2.0
Mirrord Operatormetalbear-co/mirrord5.4k1 repos~4.6kAutomated safety check: PassMIT
Syncmetapawurb/hotpath-rs1.9k—~1.2kAutomated safety check: NotesMIT
Devopsnicepkg/auto-company1952 repos~814Automated safety check: PassMIT
KubeShark for KubernetesLukasNiessen/kubernetes-skill446—~1.2kAutomated safety check: PassMIT
Optimize Slurm TopologyNVlabs/alpasim1.3k—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Mirrord Operator

    metalbear-co/mirrord

    Help users install and configure the mirrord Operator for team/enterprise environments.

    5.4k GitHub starsUsed in 1 repo~4.6k tokens
    DevOps & CloudAuto-check passed
  • Syncmeta

    pawurb/hotpath-rs

    Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta).

    1.9k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    195 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • KubeShark for Kubernetes

    LukasNiessen/kubernetes-skill

    Keeps Kubernetes manifests, Helm charts and policies grounded by diagnosing six failure modes, such as insecure defaults and API drift, and loading only matching references.

    446 GitHub stars~1.2k tokensUpdated 27 days ago
    DevOps & CloudAuto-check passed
  • Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.

    1.3k GitHub stars~1.6k tokensUpdated 23 days ago
    DevOps & CloudAuto-check passed
  • UModel Root Cause Analysis

    alibaba/UnifiedModel

    Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.

    415 GitHub stars~1.9k tokensUpdated 17 days ago
    DevOps & CloudAuto-check passed

More from google/skills

All 150 skills in this repo
  • Official

    Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.

    21k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated yesterday
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Gke AI Troubleshooting Jobset Interruption

What does Gke AI Troubleshooting Jobset Interruption do?

Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously. Gke AI Troubleshooting Jobset Interruption is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously.

When should I use Gke AI Troubleshooting Jobset Interruption?

Gke AI Troubleshooting Jobset Interruption fits situations like: troubleshooting JobSet restart loops; spot VM preemptions; Node readiness failures; coordinator worker crashes.

How do I install Gke AI Troubleshooting Jobset Interruption in Claude Code?

Run `npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-jobset-interruption in google/skills) into .claude/skills/gke-ai-troubleshooting-jobset-interruption in your project. Claude Code loads it when a task matches its description.

How do I install Gke AI Troubleshooting Jobset Interruption in Codex?

Run `npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-jobset-interruption in google/skills) into .agents/skills/gke-ai-troubleshooting-jobset-interruption in your project. Codex loads it when a task matches its description.

Can I use Gke AI Troubleshooting Jobset Interruption in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-jobset-interruption, .gemini/skills/gke-ai-troubleshooting-jobset-interruption, .github/skills/gke-ai-troubleshooting-jobset-interruption and .opencode/skills/gke-ai-troubleshooting-jobset-interruption in your project.

What does Gke AI Troubleshooting Jobset Interruption need to run?

Going by SKILL.md and its folder, Gke AI Troubleshooting Jobset Interruption needs a shell for the scripts in its folder. Our summary lists: A Bash shell.

Does Gke AI Troubleshooting Jobset Interruption access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gke AI Troubleshooting Jobset Interruption safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gke AI Troubleshooting Jobset Interruption use?

Gke AI Troubleshooting Jobset Interruption is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gke AI Troubleshooting Jobset Interruption use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 653 tokens, read only when the agent opens those files.

What are the alternatives to Gke AI Troubleshooting Jobset Interruption?

Skills that share tags, products or a category with Gke AI Troubleshooting Jobset Interruption: Mirrord Operator (metalbear-co/mirrord, 5.4k stars), Syncmeta (pawurb/hotpath-rs, 1.9k stars), Devops (nicepkg/auto-company, 195 stars) and KubeShark for Kubernetes (LukasNiessen/kubernetes-skill, 446 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gke AI Troubleshooting Jobset Interruption?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.