Mirrord Operator
metalbear-co/mirrord
Help users install and configure the mirrord Operator for team/enterprise environments.
Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-dynamic-slices-monitoring --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring .claude/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring into .claude/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-dynamic-slices-monitoring --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring .agents/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring into .agents/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-dynamic-slices-monitoring --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring .cursor/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring into .cursor/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-dynamic-slices-monitoring --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring .gemini/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring into .gemini/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-dynamic-slices-monitoringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring .github/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring into .github/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-dynamic-slices-monitoring --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring .opencode/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring into .opencode/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-dynamic-slices-monitoring", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gke-ai-troubleshooting-tpu-dynamic-slices-monitoringMonitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources.
Gke AI Troubleshooting Tpu Dynamic Slices Monitoring is an agent skill from google/skills, published by the product's own GitHub organization. Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources. Use when checking TPU slice lifecycle states, troubleshooting slice provisioning failures, validating single-slice or multi-slice (JobSet) workload manifests, or safely patching stuck finalizers and disabling the slice controller. Don't use for generic GKE cluster node pool creation or standard non-TPU workload management (use gke-basics or gke-cluster-creation instead).
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/failure_signatures.md` and `scripts/validate_queries.sh`).
It sits in DevOps & Cloud. It works with Google Kubernetes Engine. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4b940dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
kubectlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gke AI Troubleshooting Tpu Dynamic Slices Monitoring loads about 2k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 126 tokens; SKILL.md has 794 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from google/skills at commit 4b940dd, republished under its Apache-2.0 licence (© google). 794 words, ~1,958 tokens.
.claude/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Monitors the status of TPU Slice custom resources, troubleshoots provisioning failures, validates workload manifests on dynamic slices, and performs cleanups.
kubectl and gcloud CLIs configured to access the GKE cluster.Gather project, cluster, and slice context using cluster tools or the following parameters:
{project_id} (e.g., my-gcp-project){cluster_name} (e.g., tpu-cluster){location} (e.g., us-central1-a){slice_name} (e.g., test-slice){timestamp} (Optional; default to the last 30 minutes
window [T - 30m] to [T + 30m])When asked to inspect, troubleshoot, or check a slice status, immediately execute kubectl describe slice {slice_name} using available cluster tools to perform the inspection. Parse the resulting Status.Conditions output against the condition table below to diagnose the exact state and provide concrete recommendations.
Command:
kubectl describe slice {slice_name}Analyze the Status.Conditions (especially Type: Ready and its Reason and
Status):
| Lifecycle State / Reason | Meaning | Recommended Action |
|---|---|---|
SliceNotCreated | GKE Slice Controller is initializing the slice and performing resource checks. | Wait a few minutes and re-check slice status. |
SliceCreationFailed | Prerequisites validation failed (e.g., selected nodes don't exist, nodes are already used by another slice, or the topology doesn't match the number of partitions). | Verify selected nodes exist, are unallocated, and topology matches partition count. |
ACTIVATING | GKE is actively forming and provisioning the TPU slice. | Monitor node provisioning. |
ACTIVE | The TPU slice is successfully formed and ready to host workloads. | Proceed to deploy or check workloads. |
ACTIVE_DEGRADED | The slice is usable, but one or more sub-blocks are degraded. | Monitor workload logs for interconnect or device errors. Check faulty node VMs. |
FAILED | GKE failed to form the TPU slice (e.g., selected nodes are not part of the same reservation block). | Ensure all selected nodes belong to the same reservation block. |
DEACTIVATING | The slice is dismantling (triggered by user deletion or a critical systemic failure). | Wait for dismantling to finish, or patch finalizers if stuck. |
INCOMPLETE | The terminal phase before the Slice CR is deleted from the cluster. | No action required; the resource will be removed shortly. |
When investigating slice creation or provisioning failures (SliceCreationFailed or FAILED), perform the following verification steps:
kubectl get nodes -l cloud.google.com/gke-tpu-slice, kubectl get slice -A).2x2 requires 4 nodes).Ensure workload manifests are configured correctly to target the dynamic slice.
Check that the Pod template contains the following annotations and selectors:
cloud.google.com/gke-tpu-slice-topology: "{topology}" (e.g.,
"4x4x4")cloud.google.com/gke-tpu-topology: "{topology}" (e.g., "4x4x4")cloud.google.com/gke-tpu-accelerator: "{accelerator_type}" (e.g.,
"tpu7x")cloud.google.com/gke-tpu-slice: "{slice_name}" (e.g., "test-slice")If deploying a multi-slice JobSet, verify:
alpha.jobset.sigs.k8s.io/exclusive-topology: cloud.google.com/gke-tpu-slicecloud.google.com/gke-tpu-slice-topology: "{topology}"cloud.google.com/gke-tpu-topology: "{topology}"cloud.google.com/gke-tpu-accelerator: "{accelerator_type}"cloud.google.com/gke-tpu-slice in the
nodeSelector; JobSet handles slice assignment automatically.If a slice is stuck in DEACTIVATING or deletion hangs indefinitely due to stuck finalizers:
Identify Cause: Explain that finalizers on the slice resource (metadata.finalizers) are preventing Kubernetes from completing resource deletion.
Propose Resolution: Propose removing finalizers from the metadata path (/metadata/finalizers) using a JSON patch operation:
kubectl patch slice {slice_name} --type json -p='[{"op": "remove", "path": "/metadata/finalizers"}]'Provide Warning: Explicitly warn the user that removing finalizers bypasses standard controller dismantling and may leave underlying VM, network, or accelerator resources uncleaned or orphaned.
CRITICAL SAFETY MANDATE: The response MUST explicitly ask the user for confirmation (e.g. "Removing finalizers on /metadata/finalizers via JSON patch is a high-risk operation that may leave orphaned resources. Do you confirm you want to apply this patch to slice {slice_name}?") and pause for user confirmation before applying or executing the patch.
If dynamic slicing needs to be disabled:
Check for existing Slices:
kubectl get slice -AEnsure all slices are deleted before disabling the controller.
Disable Slice Controller via gcloud:
gcloud container clusters update {cluster_name} \
--location={location} \
--no-enable-slice-controllerDelete the Slice CRD:
kubectl delete crd slices.accelerator.gke.ioClean up Node Labels: Remove GKE TPU Slice labels from all nodes in the cluster:
kubectl label nodes --all cloud.google.com/gke-tpu-slice- cloud.google.com/gke-tpu-slice-topology-© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts, references) in skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring of google/skills.
Open the folder on GitHubat commit 4b940dd
Gke AI Troubleshooting Tpu Dynamic Slices Monitoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gke AI Troubleshooting Tpu Dynamic Slices Monitoring this skillgoogle/skills | 21k | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Mirrord Operatormetalbear-co/mirrord | 5.4k | 1 repos | ~4.6k | Automated safety check: Pass | MIT | |
| Devopsnicepkg/auto-company | 195 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| KubeShark for KubernetesLukasNiessen/kubernetes-skill | 446 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Kcli Cluster Deploymentkarmab/kcli | 653 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Aicr Analyzing SnapshotsNVIDIA/aicr | 440 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 |
metalbear-co/mirrord
Help users install and configure the mirrord Operator for team/enterprise environments.
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
LukasNiessen/kubernetes-skill
Keeps Kubernetes manifests, Helm charts and policies grounded by diagnosing six failure modes, such as insecure defaults and API drift, and loading only matching references.
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
NVIDIA/aicr
A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…
pifferologo/cloud-agents-cli
This skill should be used when the user wants to "deploy an agent", "deploy my ADK agent", "set up CI/CD", "configure secrets", "troubleshoot a deployment", or needs guidance on Agent Runtime, Cloud…
google/skills
Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
Works with
Categories
Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources. Gke AI Troubleshooting Tpu Dynamic Slices Monitoring is an agent skill from google/skills, published by the product's own GitHub organization. Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources.
Gke AI Troubleshooting Tpu Dynamic Slices Monitoring fits situations like: checking TPU slice lifecycle states; troubleshooting slice provisioning failures; validating single-slice; multi-slice (JobSet) workload manifests.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring in google/skills) into .claude/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring in google/skills) into .agents/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring, .gemini/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring, .github/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring and .opencode/skills/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring in your project.
Going by SKILL.md and its folder, Gke AI Troubleshooting Tpu Dynamic Slices Monitoring needs a shell for the scripts in its folder and the command-line tools its instructions call (kubectl). Our summary lists: A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Gke AI Troubleshooting Tpu Dynamic Slices Monitoring is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 406 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Gke AI Troubleshooting Tpu Dynamic Slices Monitoring: Mirrord Operator (metalbear-co/mirrord, 5.4k stars), Devops (nicepkg/auto-company, 195 stars), KubeShark for Kubernetes (LukasNiessen/kubernetes-skill, 446 stars) and Kcli Cluster Deployment (karmab/kcli, 653 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.