Signoz
qjoly/GitOps
Manage the self-hosted SigNoz observability stack in this GitOps repo.
Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-metrics-monitoring --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring .claude/skills/gke-ai-troubleshooting-tpu-metrics-monitoring && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gke-ai-troubleshooting-tpu-metrics-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring into .claude/skills/gke-ai-troubleshooting-tpu-metrics-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-metrics-monitoring", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-metrics-monitoring --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring .agents/skills/gke-ai-troubleshooting-tpu-metrics-monitoring && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gke-ai-troubleshooting-tpu-metrics-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring into .agents/skills/gke-ai-troubleshooting-tpu-metrics-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-metrics-monitoring", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-metrics-monitoring --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring .cursor/skills/gke-ai-troubleshooting-tpu-metrics-monitoring && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gke-ai-troubleshooting-tpu-metrics-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring into .cursor/skills/gke-ai-troubleshooting-tpu-metrics-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-metrics-monitoring", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-metrics-monitoring --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring .gemini/skills/gke-ai-troubleshooting-tpu-metrics-monitoring && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-metrics-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring into .gemini/skills/gke-ai-troubleshooting-tpu-metrics-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-metrics-monitoring", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-metrics-monitoringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring .github/skills/gke-ai-troubleshooting-tpu-metrics-monitoring && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-metrics-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring into .github/skills/gke-ai-troubleshooting-tpu-metrics-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-metrics-monitoring", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-metrics-monitoring --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring .opencode/skills/gke-ai-troubleshooting-tpu-metrics-monitoring && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-metrics-monitoring" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring into .opencode/skills/gke-ai-troubleshooting-tpu-metrics-monitoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-metrics-monitoring", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gke-ai-troubleshooting-tpu-metrics-monitoringMonitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL.
Gke AI Troubleshooting Tpu Metrics Monitoring is an agent skill from google/skills, published by the product's own GitHub organization. Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating MTTR or MTBI metrics for GKE TPUs. Don't use for general non-TPU GKE workload monitoring or non-metric TPU debugging.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/failure_signatures.md` and `scripts/validate_queries.sh`).
It sits in DevOps & Cloud, covering Monitoring and alerting. It works with Google Kubernetes Engine, Prometheus and Kubernetes. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 4b940dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gke AI Troubleshooting Tpu Metrics Monitoring loads about 1.9k tokens when it runs, and up to ~2.2k if it reads all its reference files. Until then it costs about 110 tokens; SKILL.md has 546 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from google/skills at commit 4b940dd, republished under its Apache-2.0 licence (© google). 546 words, ~1,860 tokens.
.claude/skills/gke-ai-troubleshooting-tpu-metrics-monitoring/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.This skill enables the agent to monitor GKE TPU workloads, nodes, and node pools using GKE system metrics. It helps diagnose if workload interruptions or performance issues are caused by underlying infrastructure.
Independently gather required context (such as cluster details or node pool names) using available GKE and Cloud tools, or use the provided {variable} placeholders:
{project_id}: The GCP Project ID.{cluster_name}: The GKE Cluster Name.{location}: The GKE Cluster Location (region or zone).{node_name}: (Optional) The name of the specific GKE node.{node_pool_name}: (Optional) The name of the GKE node pool.Before analyzing runtime metrics, verify that the workload is configured to export them. This ensures the cluster and container environment are set up for automated metric scraping and visibility into accelerator health.
containerPort: 8431 exposed on the TPU container (required for Prometheus metric scraping).0.4.14 or later if using JAX (earlier versions do not export runtime metrics).1.27.4-gke.900 or later (required for TPU runtime metric support).If configured correctly, the following metrics are available in Cloud Monitoring (monitored resources k8s_node and k8s_container):
kubernetes.io/container/accelerator/duty_cycle: Percentage of time over the past sampling period (60 seconds) during which the TensorCores were actively processing on a TPU chip.kubernetes.io/container/accelerator/memory_used: Amount of accelerator memory allocated in bytes.kubernetes.io/container/accelerator/memory_total: Total accelerator memory in bytes.kubernetes.io/node/accelerator/duty_cyclekubernetes.io/node/accelerator/memory_usedkubernetes.io/node/accelerator/memory_totalQuery the status condition of GKE nodes (GKE version 1.32.1-gke.1357001 or later).
kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", node_name="{node_name}", condition="Ready", status="True"}kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", condition!="Ready", status="True"}kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", condition="Ready", status="False"}avg by (condition,status)(avg_over_time(kubernetes_io:node_status_condition{monitored_resource="k8s_node"}[5m]))Query the status of multi-host TPU node pools.
kubernetes_io:node_pool_status{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}", node_pool_name="{node_pool_name}", status="Running"}count by (status)(count_over_time(kubernetes_io:node_pool_status{monitored_resource="k8s_node_pool"}[5m]))Provisioning, Running, Error, Reconciling, Stopping.Query if all nodes in a multi-host TPU node pool are available.
avg by (node_pool_name)(avg_over_time(kubernetes_io:node_pool_multi_host_available{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}"}[5m]))1 (True, all nodes available) or 0 (False, some nodes unavailable).Query the count of interruptions for GKE nodes.
sum by (interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node"}[5m]))TerminationEvent, MaintenanceEvent, PreemptionEvent.
Interruption Reasons: HostError, Eviction, AutoRepair.sum by (interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node", interruption_reason="HW/SW Maintenance"}[5m]))sum by (node_pool_name,interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_pool_interruption_count{monitored_resource="k8s_node_pool", interruption_reason="HW/SW Maintenance", node_pool_name="{node_pool_name}"}[5m]))Calculate Mean Time to Recovery (MTTR) and Mean Time Between Interruptions (MTBI) over the last 7 days.
sum(sum_over_time(kubernetes_io:node_pool_accelerator_times_to_recover_sum{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}"}[7d])) / sum(sum_over_time(kubernetes_io:node_pool_accelerator_times_to_recover_count{monitored_resource="k8s_node_pool",cluster_name="{cluster_name}"}[7d]))sum(count_over_time(kubernetes_io:node_memory_total_bytes{monitored_resource="k8s_node", node_name=~"gke-tpu.*|gk3-tpu.*", cluster_name="{cluster_name}"}[7d])) / sum(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node", node_name=~"gke-tpu.*|gk3-tpu.*", cluster_name="{cluster_name}"}[7d]))For GKE version 1.28.1-gke.1066000 or later, monitor TPU host performance.
kubernetes.io/container/accelerator/tensorcore_utilization: Current percentage of the TensorCore that is utilized.kubernetes.io/container/accelerator/memory_bandwidth_utilization: Current percentage of the accelerator memory bandwidth that is being used.kubernetes.io/node/accelerator/tensorcore_utilizationkubernetes.io/node/accelerator/memory_bandwidth_utilization© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts, references) in skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring of google/skills.
Open the folder on GitHubat commit 4b940dd
Gke AI Troubleshooting Tpu Metrics Monitoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gke AI Troubleshooting Tpu Metrics Monitoring this skillgoogle/skills | 21k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Signozqjoly/GitOps | 112 | — | ~6.1k | Automated safety check: Pass | WTFPL | |
| Proto Backend Moduleaide-family/moon | 253 | — | ~4.1k | Automated safety check: Pass | None | |
| Prometheus GrafanaBagelHole/DevOps-Security-Agent-Skills | 1.2k | — | ~2.5k | Automated safety check: Pass | MIT | |
| Grafana Dashboardpando85/kaniop | 132 | — | ~987 | Automated safety check: Pass | AGPL-3.0 | |
| Alloygrafana/skills | 282 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 |
qjoly/GitOps
Manage the self-hosted SigNoz observability stack in this GitOps repo.
aide-family/moon
Implements backend modules from proto definitions for goddess, marksman, and rabbit apps.
BagelHole/DevOps-Security-Agent-Skills
Set up metrics collection and visualization with Prometheus and Grafana.
pando85/kaniop
Improve and validate the Kaniop Grafana dashboard against repository metrics and the grigri live cluster.
grafana/skills
Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /…
qdrant/skills
Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization.
google/skills
Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
Categories
Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Gke AI Troubleshooting Tpu Metrics Monitoring is an agent skill from google/skills, published by the product's own GitHub organization. Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL.
Gke AI Troubleshooting Tpu Metrics Monitoring fits situations like: monitoring TensorCore duty cycle; multi-host TPU node pool availability; host maintenance; preemption interruptions.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring in google/skills) into .claude/skills/gke-ai-troubleshooting-tpu-metrics-monitoring in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring in google/skills) into .agents/skills/gke-ai-troubleshooting-tpu-metrics-monitoring in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-metrics-monitoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-tpu-metrics-monitoring, .gemini/skills/gke-ai-troubleshooting-tpu-metrics-monitoring, .github/skills/gke-ai-troubleshooting-tpu-metrics-monitoring and .opencode/skills/gke-ai-troubleshooting-tpu-metrics-monitoring in your project.
Going by SKILL.md and its folder, Gke AI Troubleshooting Tpu Metrics Monitoring needs a shell for the scripts in its folder. Our summary lists: A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Gke AI Troubleshooting Tpu Metrics Monitoring is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 361 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Gke AI Troubleshooting Tpu Metrics Monitoring: Signoz (qjoly/GitOps, 112 stars), Proto Backend Module (aide-family/moon, 253 stars), Prometheus Grafana (BagelHole/DevOps-Security-Agent-Skills, 1.2k stars) and Grafana Dashboard (pando85/kaniop, 132 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.