Devops
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
Diagnose GKE Cloud TPU multi-slice training hangs (Megascale HANGDETECTED logs and hang monitored events) using the ML Diagnostics Megascale XLA (MXLA) Hang Analyzer (gcloud alpha mldiagnostics…
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-mxla-hang -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-mxla-hang --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang .claude/skills/gke-ai-troubleshooting-tpu-mxla-hang && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gke-ai-troubleshooting-tpu-mxla-hang" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang into .claude/skills/gke-ai-troubleshooting-tpu-mxla-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-mxla-hang", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hangType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-mxla-hang -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-mxla-hang --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang .agents/skills/gke-ai-troubleshooting-tpu-mxla-hang && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gke-ai-troubleshooting-tpu-mxla-hang" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang into .agents/skills/gke-ai-troubleshooting-tpu-mxla-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-mxla-hang", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-mxla-hang -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-mxla-hang --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang .cursor/skills/gke-ai-troubleshooting-tpu-mxla-hang && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gke-ai-troubleshooting-tpu-mxla-hang" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang into .cursor/skills/gke-ai-troubleshooting-tpu-mxla-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-mxla-hang", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-mxla-hang -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-mxla-hang --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang .gemini/skills/gke-ai-troubleshooting-tpu-mxla-hang && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-mxla-hang" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang into .gemini/skills/gke-ai-troubleshooting-tpu-mxla-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-mxla-hang", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-mxla-hangInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-mxla-hang -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang .github/skills/gke-ai-troubleshooting-tpu-mxla-hang && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-mxla-hang" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang into .github/skills/gke-ai-troubleshooting-tpu-mxla-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-mxla-hang", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-mxla-hang -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-mxla-hang --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang .opencode/skills/gke-ai-troubleshooting-tpu-mxla-hang && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-mxla-hang" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang into .opencode/skills/gke-ai-troubleshooting-tpu-mxla-hang/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-mxla-hang", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gke-ai-troubleshooting-tpu-mxla-hangDiagnose GKE Cloud TPU multi-slice training hangs (Megascale HANGDETECTED logs and hang monitored events) using the ML Diagnostics Megascale XLA (MXLA) Hang Analyzer (gcloud alpha mldiagnostics…
Gke AI Troubleshooting Tpu Mxla Hang is an agent skill from google/skills, published by the product's own GitHub organization. Diagnose GKE Cloud TPU multi-slice training hangs (Megascale HANGDETECTED logs and hang monitored events) using the ML Diagnostics Megascale XLA (MXLA) Hang Analyzer (gcloud alpha mldiagnostics monitored-events) and 1-minute Cloud Monitoring multi-slice latency metrics (kubernetes.io/container/multislice/). Distinguishes XLA compiler/HLO launch divergence and host data-input stalls from TPU chip, SparseCore, ICI, or network fabric faults. Use when multi-slice TPU training jobs freeze without progressing steps…
Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/path-a-compiler-hlo-divergence.md`, `references/path-b-program-queueing-input-stall.md` and `references/path-c-hardware-network-faults.md`).
It sits in DevOps & Cloud, covering Container orchestration. It works with Google Kubernetes Engine, Google Cloud and Kubernetes. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4b940dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gcloudkubectlFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.cloud.google.comcloud.google.comopenxla.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gke AI Troubleshooting Tpu Mxla Hang loads about 3.7k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 206 tokens; SKILL.md has 1,313 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from google/skills at commit 4b940dd, republished under its Apache-2.0 licence (© google). 1,313 words, ~3,706 tokens.
.claude/skills/gke-ai-troubleshooting-tpu-mxla-hang/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Diagnose Cloud TPU multi-slice training hangs on Google Kubernetes Engine (GKE)
by correlating Megascale HANG_DETECTED logs and ML Diagnostics Workload
Monitoring Megascale XLA (MXLA) Hang Analyzer reports with 1-minute
multi-slice latency metrics (kubernetes.io/container/multislice/*) and GKE
node topology labels.
gcloud with
alpha component for gcloud alpha mldiagnostics) and kubectl.gcloud billing projects describe {project_id}), authenticate
(gcloud auth login), set the target project (gcloud config set project {project_id}), and ensure container.googleapis.com,
logging.googleapis.com, monitoring.googleapis.com, and
hypercomputecluster.googleapis.com are enabled.jobset and job GKE
job types, and is compatible with GKE versions 1.36.0-gke.4681000 and later
(Configure GKE for ML Diagnostics,
which also covers the cluster setup needed for on-demand profiling in Step 4
Path B). If a workload uses another framework (such as PyTorch) or another
custom resource type, gcloud alpha mldiagnostics won't list ML runs or
monitored events for it. The Megascale XLA hang analyzer and Megascale XLA
metrics require LibTPU 0.40.0 or later (see "Get started" in
Workload monitoring with ML Diagnostics).roles/hypercomputecluster.editor), the role
listed in the "IAM permissions" section of
ML Diagnostics platform,
for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored
events, and on-demand profiler sessions)roles/monitoring.viewer) for the PromQL queries in Step
2roles/logging.viewer) for the Cloud Logging query in Step 1roles/container.viewer) for the kubectl get nodes query in Step 3[High Risk] steps): Kubernetes Engine Cluster Admin
(roles/container.clusterAdmin)Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.
When you recommend a fix, link the doc section that describes it.
A Megascale hang occurs when a multi-slice worker has waited on a Megascale
communication operation for a set timeout period. The TPU logs then show a
Megascale HANG_DETECTED message. HANG_DETECTED is a catch-all signal that
the workload isn't progressing, and the cause can be in software or in hardware.
When a hang occurs, ML Diagnostics runs the Megascale XLA (MXLA) Hang Analyzer, which reports the likely cause as a code. Consult the
Megascale XLA (MXLA) Hang Analyzer
section for the definition and recommended action of each code, and route by
category:
FINGERPRINT_MISMATCH), route to
Path A: Compiler or HLO divergence.DATA_INPUT_STALL), route to
Path B: Host program queueing or data input stall.NOT_DETECTED
or reports UNKNOWN:HANG_DETECTED logs or a hang monitored event fired, but the analyzer
report has detectionState: "NOT_DETECTED" or reports UNKNOWN ("The MXLA
hang was detected but a potential cause is not determined"), route to
Path D: No DETECTED analyzer or UNKNOWN cause.[Low Risk]Collect the target parameters. By default, query a 60-minute window `[T - 30m, T
around{issue_time}`:{project_id}: Google Cloud project ID{location}: Google Cloud region where the ML run and GKE cluster reside (for
example, us-central1){cluster_name}: GKE cluster name{namespace} / {workload_name}: Kubernetes namespace and JobSet/Pod prefix{ml_run_id}: ML Diagnostics run ID{issue_time}: Timestamp when the hang occurred (T, ISO-8601 UTC){start_time}: T - 30m{end_time}: T + 30mHANG_DETECTED logs and hang events [Low Risk]HANG_DETECTED (read-only Cloud Logging
LQL): Query k8s_container logs over [{start_time}, {end_time}] to
confirm HANG_DETECTED and identify the first stalled pods:resource.type="k8s_container"
resource.labels.project_id="{project_id}"
resource.labels.cluster_name="{cluster_name}"
"HANG_DETECTED"
timestamp >= "{start_time}" AND timestamp <= "{end_time}"gcloud alpha mldiagnostics machine-learning-run list) to identify {ml_run_id}.gcloud alpha mldiagnostics monitored-events list and gcloud alpha mldiagnostics monitored-events describe) or
Access Workload Monitoring information through the API
to inspect the Megascale XLA (MXLA) Hang Analyzer report (detectionState,
details, and recommendedActions).HANG_DETECTED logs or a hang monitored event fired, run Step 2 to
correlate with the 1-minute metrics; if no analyzer reports
detectionState: "DETECTED" or the analyzer reports UNKNOWN, follow
Path D: No DETECTED analyzer or UNKNOWN cause.HANG_DETECTED log or hang monitored event exists and multi-slice
latencies are normal in Step 2, rule out an MXLA hang by following
Path E: Healthy telemetry.[Low Risk]Consult the
System Metrics
section of the Workload Monitoring guide for the 1-minute multi-slice network
(kubernetes.io/container/multislice/network/*), multi-slice accelerator
(kubernetes.io/container/multislice/accelerator/*), and node duty-cycle
(kubernetes.io/node/accelerator/duty_cycle) metrics, and run read-only PromQL
queries over [{start_time}, {end_time}]:
# 1. P95 multi-slice collective end-to-end latency by pod
histogram_quantile(
0.95,
sum by (pod_name, le) (
rate(kubernetes_io:container_multislice_network_collective_end_to_end_latencies_bucket{
monitored_resource="k8s_container",
project_id="{project_id}",
cluster_name="{cluster_name}"
}[5m])
)
)
# 2. P95 multi-slice DCN transfer latency by pod
histogram_quantile(
0.95,
sum by (pod_name, le) (
rate(kubernetes_io:container_multislice_network_dcn_transfer_latencies_bucket{
monitored_resource="k8s_container",
project_id="{project_id}",
cluster_name="{cluster_name}"
}[5m])
)
)
# 3. P95 host-to-device transfer latency by pod
histogram_quantile(
0.95,
sum by (pod_name, le) (
rate(kubernetes_io:container_multislice_accelerator_host_to_device_transfer_latencies_bucket{
monitored_resource="k8s_container",
project_id="{project_id}",
cluster_name="{cluster_name}"
}[5m])
)
)
# 4. Node TPU duty cycle (Workload Monitoring detects a hang as a prolonged
# period of minimal to no TPU activity)
kubernetes_io:node_accelerator_duty_cycle{
monitored_resource="k8s_node",
project_id="{project_id}",
cluster_name="{cluster_name}"
}[Low Risk]When the Megascale XLA (MXLA) Hang Analyzer reports culprit numeric Compute
Engine instance IDs in details or recommendedActions, map those numeric IDs
to GKE Node names and physical topology blocks using this read-only kubectl
query inspecting container.googleapis.com/instance_id:
kubectl get nodes -l cloud.google.com/gke-tpu-accelerator \
-o jsonpath='{range .items[*]}{.metadata.name}{"\tinstance_id="}{.metadata.annotations.container\.googleapis\.com/instance_id}{"\tblock="}{.metadata.labels.cloud\.google\.com/gce-topology-block}{"\tsubblock="}{.metadata.labels.cloud\.google\.com/gce-topology-subblock}{"\thost="}{.metadata.labels.cloud\.google\.com/gce-topology-host}{"\n"}{end}'Load and follow only the reference file that matches the root-cause category from Step 1:
FINGERPRINT_MISMATCH): Read
Path A: Compiler or HLO divergence.PROGRAM_NOT_QUEUED or
DATA_INPUT_STALL): Read
Path B: Host program queueing or data input stall.UNRECOVERABLE_ERROR on specific instances): Read
Path C: Hardware or network faults.HANG_DETECTED logs or hang event fired, but the analyzer is
NOT_DETECTED or reports UNKNOWN): Read
Path D: No DETECTED analyzer or UNKNOWN cause.HANG_DETECTED logs or hang events, and steady telemetry):
Read Path E: Healthy telemetry.gcloud compute instance-groups managed commands, such as
delete, on a node pool's managed instance group. Handle nodes through GKE,
as described in
Path C: Hardware or network faults.© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang of google/skills.
Open the folder on GitHubat commit 4b940dd
Gke AI Troubleshooting Tpu Mxla Hang next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gke AI Troubleshooting Tpu Mxla Hang this skillgoogle/skills | 21k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | |
| Devopsnicepkg/auto-company | 195 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| Kcli Cluster Deploymentkarmab/kcli | 653 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| GCP Gkesickn33/agentic-awesome-skills | 47k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Apex Azure Cloud Migratejonathan-vella/apex | 217 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Dt Obs GCPDynatrace/dynatrace-for-ai | 163 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 |
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
sickn33/agentic-awesome-skills
Deploy and manage Google Kubernetes Engine clusters. An agent skill from sickn33/agentic-awesome-skills.
jonathan-vella/apex
WORKFLOW SKILL — Assess and migrate cross-cloud workloads to Azure: assessments and code conversion from AWS, GCP, Heroku, Kubernetes or Spring.
Dynatrace/dynatrace-for-ai
GCP cloud resources including Compute Engine, GKE, Cloud Run, Pub/Sub, VPC networking, DNS, IAM, Secret Manager, and monitoring.
metalbear-co/mirrord
Help users install and configure the mirrord Operator for team/enterprise environments.
google/skills
Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
Categories
Diagnose GKE Cloud TPU multi-slice training hangs (Megascale HANGDETECTED logs and hang monitored events) using the ML Diagnostics Megascale XLA (MXLA) Hang Analyzer (gcloud alpha mldiagnostics…. Gke AI Troubleshooting Tpu Mxla Hang is an agent skill from google/skills, published by the product's own GitHub organization.io/container/multislice/).
Gke AI Troubleshooting Tpu Mxla Hang fits situations like: multi-slice TPU training jobs freeze without progressing steps; emit HANGDETECTED; stall in collective operations; gradual step-time throughput drops without hangs (use gke-ai-troubleshooting-tpu-performance-degradation).
Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-mxla-hang -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang in google/skills) into .claude/skills/gke-ai-troubleshooting-tpu-mxla-hang in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-mxla-hang -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang in google/skills) into .agents/skills/gke-ai-troubleshooting-tpu-mxla-hang in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-mxla-hang -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-tpu-mxla-hang, .gemini/skills/gke-ai-troubleshooting-tpu-mxla-hang, .github/skills/gke-ai-troubleshooting-tpu-mxla-hang and .opencode/skills/gke-ai-troubleshooting-tpu-mxla-hang in your project.
Going by SKILL.md and its folder, Gke AI Troubleshooting Tpu Mxla Hang needs the command-line tools its instructions call (gcloud and kubectl).
SKILL.md names 3 domains. As links in the text: docs.cloud.google.com, cloud.google.com and openxla.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gke AI Troubleshooting Tpu Mxla Hang is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Gke AI Troubleshooting Tpu Mxla Hang: Devops (nicepkg/auto-company, 195 stars), Kcli Cluster Deployment (karmab/kcli, 653 stars), GCP Gke (sickn33/agentic-awesome-skills, 47k stars) and Apex Azure Cloud Migrate (jonathan-vella/apex, 217 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.