Mirrord Operator
metalbear-co/mirrord
Help users install and configure the mirrord Operator for team/enterprise environments.
Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE.
$ npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-handle-disruption-gpu-tpu --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu .claude/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gke-ai-troubleshooting-handle-disruption-gpu-tpu" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu into .claude/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-handle-disruption-gpu-tpu", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpuType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-handle-disruption-gpu-tpu --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu .agents/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gke-ai-troubleshooting-handle-disruption-gpu-tpu" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu into .agents/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-handle-disruption-gpu-tpu", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-handle-disruption-gpu-tpu --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu .cursor/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gke-ai-troubleshooting-handle-disruption-gpu-tpu" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu into .cursor/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-handle-disruption-gpu-tpu", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-handle-disruption-gpu-tpu --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu .gemini/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-handle-disruption-gpu-tpu" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu into .gemini/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-handle-disruption-gpu-tpu", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills gke-ai-troubleshooting-handle-disruption-gpu-tpuInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu .github/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-handle-disruption-gpu-tpu" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu into .github/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-handle-disruption-gpu-tpu", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-handle-disruption-gpu-tpu --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu .opencode/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-handle-disruption-gpu-tpu" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu into .opencode/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-handle-disruption-gpu-tpu", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gke-ai-troubleshooting-handle-disruption-gpu-tpuDiagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE.
Gke AI Troubleshooting Handle Disruption GPU Tpu is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE. Use when diagnosing node disruptions, predicting host maintenance events on GPU/TPU nodepools, inspecting node interruption PromQL metrics, auditing node taints, or configuring workload protection strategies (graceful termination, opportunistic maintenance, PodDisruptionBudgets). Don't use for general GKE cluster creation, network policy…
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud. It works with Google Kubernetes Engine, Prometheus and Kubernetes. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4b940dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.cloud.google.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gke AI Troubleshooting Handle Disruption GPU Tpu loads about 1.6k tokens when it runs. Until then it costs about 153 tokens; SKILL.md has 601 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from google/skills at commit 4b940dd, republished under its Apache-2.0 licence (© google). 601 words, ~1,631 tokens.
.claude/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu/SKILL.md (or your agent's skills folder).project_id, location, cluster_name, timestamp) BEFORE
delivering theories or general diagnostic commands. Only skip context
acquisition if the user explicitly requests a generic reusable runbook or
provides a complete static telemetry/log dump for offline analysis.node_name, workload_name, workload_namespace,
nodepool_name.Action: Propose running kubectl to check if nodes have the scheduled
maintenance label indicating an upcoming disruption.
Example Command:
kubectl get nodes -l cloud.google.com/scheduled-maintenance-time -L cloud.google.com/scheduled-maintenance-timeInterpretation: The SCHEDULED-MAINTENANCE-TIME column shows the Unix
epoch time when the VM is scheduled for maintenance. If this label exists, a
disruption is guaranteed to occur.
Action: Call any available monitoring tool or provide PromQL for manual verification.
Mandatory Monitoring Rule: Whenever recommending follow-up monitoring or
interruption tracking over time, you MUST explicitly present a PromQL
query using the metric kubernetes_io:node_interruption_count filtered by
interruption_reason="HW/SW Maintenance". Do not suggest general Cloud
Monitoring dashboards or Metrics Explorer without providing this specific
PromQL metric expression.
Example Query:
# Fetch host maintenance events for nodes
sum by (interruption_type,interruption_reason)( sum_over_time( kubernetes_io:node_interruption_count{monitored_resource="k8s_node", interruption_reason="HW/SW Maintenance"}[${__interval}]))# See the interruption count aggregated by node pool
sum by (node_pool_name,interruption_type,interruption_reason)( sum_over_time( kubernetes_io:node_pool_interruption_count{monitored_resource="k8s_node_pool", interruption_reason="HW/SW Maintenance", node_pool_name="{nodepool_name}" }[${__interval}]))Interpretation: If kubernetes_io:node_interruption_count shows
values > 0 for interruption_reason="HW/SW Maintenance", it indicates the
underlying Compute Engine VM was interrupted due to scheduled host
maintenance.
query_logs or instruct the user to filter their GKE logs
for active host maintenance events, and check node taints.cloud.google.com/active-node-maintenance is set to ONGOING. To check if
GKE has cordoned the terminating node to prevent new workloads from being
scheduled, verify whether the
cloud.google.com/impending-node-termination:NoSchedule taint is present
(either in GKE event logs or directly via kubectl describe node).cloud.google.com/active-node-maintenance set to ONGOING means
workloads are actively being stopped by GKE due to host maintenance.cloud.google.com/impending-node-termination:NoSchedule taint means GKE
has cordoned the node to prevent new Pods from being scheduled on the
terminating node. DO NOT recommend tolerating this taint.spec.terminationGracePeriodSeconds (up to 60 minutes) to
handle the SIGTERM signal before node shutdown.PodDisruptionBudget to maintain minAvailable replicas during
evictions and disruptions.© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu of google/skills.
Open the folder on GitHubat commit 4b940dd
Gke AI Troubleshooting Handle Disruption GPU Tpu next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gke AI Troubleshooting Handle Disruption GPU Tpu this skillgoogle/skills | 21k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Mirrord Operatormetalbear-co/mirrord | 5.4k | 1 repos | ~4.6k | Automated safety check: Pass | MIT | |
| Devopsnicepkg/auto-company | 195 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| KubeShark for KubernetesLukasNiessen/kubernetes-skill | 446 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Kcli Cluster Deploymentkarmab/kcli | 653 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Aicr Analyzing SnapshotsNVIDIA/aicr | 440 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 |
metalbear-co/mirrord
Help users install and configure the mirrord Operator for team/enterprise environments.
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
LukasNiessen/kubernetes-skill
Keeps Kubernetes manifests, Helm charts and policies grounded by diagnosing six failure modes, such as insecure defaults and API drift, and loading only matching references.
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
NVIDIA/aicr
A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…
qjoly/GitOps
Manage the self-hosted SigNoz observability stack in this GitOps repo.
google/skills
Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
Categories
Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE. Gke AI Troubleshooting Handle Disruption GPU Tpu is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE.
Gke AI Troubleshooting Handle Disruption GPU Tpu fits situations like: diagnosing node disruptions; predicting host maintenance events on GPU/TPU nodepools; inspecting node interruption PromQL metrics; auditing node taints.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu in google/skills) into .claude/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-handle-disruption-gpu-tpu in google/skills) into .agents/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpu -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu, .gemini/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu, .github/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu and .opencode/skills/gke-ai-troubleshooting-handle-disruption-gpu-tpu in your project.
Going by SKILL.md and its folder, Gke AI Troubleshooting Handle Disruption GPU Tpu needs the command-line tools its instructions call (kubectl).
SKILL.md names 1 domain. As links in the text: docs.cloud.google.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gke AI Troubleshooting Handle Disruption GPU Tpu is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Gke AI Troubleshooting Handle Disruption GPU Tpu: Mirrord Operator (metalbear-co/mirrord, 5.4k stars), Devops (nicepkg/auto-company, 195 stars), KubeShark for Kubernetes (LukasNiessen/kubernetes-skill, 446 stars) and Kcli Cluster Deployment (karmab/kcli, 653 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.