Mirrord Operator
metalbear-co/mirrord
Help users install and configure the mirrord Operator for team/enterprise environments.
Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously.
$ npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-jobset-interruption --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-jobset-interruption .claude/skills/gke-ai-troubleshooting-jobset-interruption && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gke-ai-troubleshooting-jobset-interruption" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-jobset-interruption into .claude/skills/gke-ai-troubleshooting-jobset-interruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-jobset-interruption", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-jobset-interruptionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-jobset-interruption --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-jobset-interruption .agents/skills/gke-ai-troubleshooting-jobset-interruption && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gke-ai-troubleshooting-jobset-interruption" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-jobset-interruption into .agents/skills/gke-ai-troubleshooting-jobset-interruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-jobset-interruption", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-jobset-interruption --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-jobset-interruption .cursor/skills/gke-ai-troubleshooting-jobset-interruption && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gke-ai-troubleshooting-jobset-interruption" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-jobset-interruption into .cursor/skills/gke-ai-troubleshooting-jobset-interruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-jobset-interruption", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/gke-ai-troubleshooting-jobset-interruption--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-jobset-interruption --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-jobset-interruption .gemini/skills/gke-ai-troubleshooting-jobset-interruption && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-jobset-interruption" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-jobset-interruption into .gemini/skills/gke-ai-troubleshooting-jobset-interruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-jobset-interruption", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills gke-ai-troubleshooting-jobset-interruptionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-jobset-interruption .github/skills/gke-ai-troubleshooting-jobset-interruption && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-jobset-interruption" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-jobset-interruption into .github/skills/gke-ai-troubleshooting-jobset-interruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-jobset-interruption", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-jobset-interruption --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-jobset-interruption .opencode/skills/gke-ai-troubleshooting-jobset-interruption && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-jobset-interruption" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-jobset-interruption into .opencode/skills/gke-ai-troubleshooting-jobset-interruption/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-jobset-interruption", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gke-ai-troubleshooting-jobset-interruptionDiagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously.
Gke AI Troubleshooting Jobset Interruption is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously. Use when troubleshooting JobSet restart loops, spot VM preemptions, node readiness failures, host VM issues, or coordinator worker crashes. Don't use for general GKE cluster creation, basic workload deployment, or non-JobSet application issues.
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/failure_signatures.md` and `scripts/validate_queries.sh`).
It sits in DevOps & Cloud. It works with Google Kubernetes Engine and Prometheus. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4b940dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gke AI Troubleshooting Jobset Interruption loads about 2.8k tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 98 tokens; SKILL.md has 1,003 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from google/skills at commit 4b940dd, republished under its Apache-2.0 licence (© google). 1,003 words, ~2,759 tokens.
.claude/skills/gke-ai-troubleshooting-jobset-interruption/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Use this skill to systematically diagnose and resolve JobSet interruptions, restarts, and preemptions on GKE clusters hosting large-scale AI/ML workloads.
kube-state-metrics for your
cluster.403 Permission Denied, authentication errors, or network
isolation, do NOT enter authentication or credential troubleshooting
loops. Populate the query templates with the acquired variables
({project_id}, {cluster_name}, {workload_name}, {start_time},
{end_time}), inspect any locally staged telemetry or mock data files if
available, and complete the diagnostic workflow and resolution
recommendations autonomously.Independently gather context using tools, workspace files, environment details, or user prompt context:
{project_id}){cluster_name}){workload_name}){namespace}){issue_time})If specific variables are not explicitly provided by the user, inspect cluster
resources or logs to determine them, or use the {variable} placeholders
provided.
{issue_time} is available (or
calculated as T), set {start_time} = T - 30m and {end_time} = T + 30m.Verify if the JobSet is experiencing restart loops and determine the frequency of restarts.
MQL Query Specification:
fetch prometheus_target
| metric 'prometheus.googleapis.com/kube_jobset_restarts/gauge'
| filter resource.cluster_name == '{cluster_name}' && metric.jobset_name == '{workload_name}'
| align next_older(1m)
| every 1m
| group_by [metric.jobset_name], [val: max(value)]PromQL Query Specification:
kube_jobset_restarts{jobset_name="{workload_name}", cluster="{cluster_name}"}Diagnostic Logic: A non-zero or increasing value for restarts indicates that the JobSet is being actively restarted by the controller due to worker failure or interruption.
Automation: Proceed to Step 2 automatically after reporting findings.
Determine if the JobSet restarts were triggered by physical nodepool-level events (such as spot preemptions, maintenance, or host terminations).
MQL Query Specification:
fetch k8s_node_pool
| metric 'kubernetes.io/node_pool/interruption_count'
| filter cluster_name == '{cluster_name}'
| align next_older(10m)
| every 10m
| group_by [metric.interruption_type, metric.interruption_reason, metadata.system.node_pool_name], [val: sum(value)]PromQL Query Specification:
sum by (interruption_type, interruption_reason, node_pool_name, cluster_name) (
avg_over_time(kubernetes_io:node_pool_interruption_count{cluster_name="{cluster_name}"}[10m])
)LQL Log Filter Specification:
resource.type="gke_nodepool"
AND resource.labels.cluster_name="{cluster_name}"
AND timestamp >= "{start_time}"
AND timestamp <= "{end_time}"Diagnostic Logic:
interruption_reason
or logs for host issues.Automation: Proceed to Step 3 automatically.
Correlate node readiness failures with physical host VMs to see if a single faulty host repeatedly fails coordinator pods.
MQL Query Specification:
fetch k8s_node
| metric 'kubernetes.io/node/status_condition'
| filter cluster_name == '{cluster_name}' && metric.condition == 'Ready' && metric.status == 'False'
| align next_older(1m)
| every 1m
| group_by [node_name, metadata.user.gke_nodepool], [val: max(value)]PromQL Query Specification:
sum by (status, condition, node_pool_name) (
kubernetes_io:node_status_condition{cluster_name="{cluster_name}", condition="Ready", status="False"}
)MQL Query Specification:
fetch k8s_node
| metric 'kubernetes.io/node/cpu/total_cores'
| filter cluster_name == '{cluster_name}'
| align next_older(1m)
| every 1m
| group_by [node_name, metadata.user.gce_topology_host, metadata.user.gke_nodepool], [val: max(value)]LQL Log Filter Specification:
resource.type="k8s_node"
AND resource.labels.cluster_name="{cluster_name}"
AND (textPayload:"host error" OR textPayload:"kernel panic" OR textPayload:"hardware failure" OR textPayload:"NodeNotReady")
AND timestamp >= "{start_time}"
AND timestamp <= "{end_time}"Diagnostic Logic: Identify if specific nodes are unhealthy
(Ready=False or Unknown) and correlate them to their GCE physical host
ID via metadata.user.gce_topology_host. Check if the same host is
repeatedly failing.
Automation: Proceed to Step 4 automatically.
Analyze pod status phases and retrieve coordinator worker logs to identify application-level crashes or network deadlocks.
Required Execution Order: You MUST analyze pod status phases (Section A) and unschedulable pod metrics (Section B) to assess overall workload health before inspecting specific worker container logs (Section C).
MQL Query Specification:
fetch k8s_pod
| metric 'kubernetes.io/pod/status/phase'
| filter cluster_name == '{cluster_name}' && pod_name ==~ '{workload_name}.*'
| align next_older(10m)
| every 10m
| group_by [metric.phase], [val: count()]PromQL Query Specification:
sum by (phase) (
avg_over_time(kube_pod_status_phase{cluster="{cluster_name}", pod=~"{workload_name}.*"}[10m])
)MQL Query Specification:
fetch k8s_pod
| metric 'kubernetes.io/pod/status/unschedulable'
| filter cluster_name == '{cluster_name}' && pod_name ==~ '{workload_name}.*'
| align next_older(10m)
| every 10m
| group_by [pod_name], [val: max(value)]LQL Log Filter Specification:
resource.type="k8s_container"
AND resource.labels.cluster_name="{cluster_name}"
AND labels."k8s-pod/jobset_sigs_k8s_io/jobset-name"="{workload_name}"
AND timestamp >= "{start_time}"
AND timestamp <= "{end_time}"Diagnostic Logic:
Automation: Proceed to Resolution.
If Step 2 showed high preemption counts on Spot VMs:
If Step 3 identified a specific host ID (gce-topology-host) that consistently
fails or triggers restarts across multiple attempts:
{start_time} ({issue_time} - 30m) and
{end_time} ({issue_time} + 30m) window.© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts, references) in skills/cloud/gke-ai-troubleshooting-jobset-interruption of google/skills.
Open the folder on GitHubat commit 4b940dd
Gke AI Troubleshooting Jobset Interruption next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gke AI Troubleshooting Jobset Interruption this skillgoogle/skills | 21k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Mirrord Operatormetalbear-co/mirrord | 5.4k | 1 repos | ~4.6k | Automated safety check: Pass | MIT | |
| Syncmetapawurb/hotpath-rs | 1.9k | — | ~1.2k | Automated safety check: Notes | MIT | |
| Devopsnicepkg/auto-company | 195 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| KubeShark for KubernetesLukasNiessen/kubernetes-skill | 446 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Optimize Slurm TopologyNVlabs/alpasim | 1.3k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 |
metalbear-co/mirrord
Help users install and configure the mirrord Operator for team/enterprise environments.
pawurb/hotpath-rs
Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta).
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
LukasNiessen/kubernetes-skill
Keeps Kubernetes manifests, Helm charts and policies grounded by diagnosing six failure modes, such as insecure defaults and API drift, and loading only matching references.
NVlabs/alpasim
Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.
alibaba/UnifiedModel
Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.
google/skills
Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
Works with
Categories
Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously. Gke AI Troubleshooting Jobset Interruption is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously.
Gke AI Troubleshooting Jobset Interruption fits situations like: troubleshooting JobSet restart loops; spot VM preemptions; Node readiness failures; coordinator worker crashes.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-jobset-interruption in google/skills) into .claude/skills/gke-ai-troubleshooting-jobset-interruption in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-jobset-interruption in google/skills) into .agents/skills/gke-ai-troubleshooting-jobset-interruption in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-jobset-interruption -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-jobset-interruption, .gemini/skills/gke-ai-troubleshooting-jobset-interruption, .github/skills/gke-ai-troubleshooting-jobset-interruption and .opencode/skills/gke-ai-troubleshooting-jobset-interruption in your project.
Going by SKILL.md and its folder, Gke AI Troubleshooting Jobset Interruption needs a shell for the scripts in its folder. Our summary lists: A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Gke AI Troubleshooting Jobset Interruption is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 653 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Gke AI Troubleshooting Jobset Interruption: Mirrord Operator (metalbear-co/mirrord, 5.4k stars), Syncmeta (pawurb/hotpath-rs, 1.9k stars), Devops (nicepkg/auto-company, 195 stars) and KubeShark for Kubernetes (LukasNiessen/kubernetes-skill, 446 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.