Devops
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (gcloud alpha mldiagnostics monitored-events /…
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-performance-degradation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-performance-degradation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation .claude/skills/gke-ai-troubleshooting-tpu-performance-degradation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gke-ai-troubleshooting-tpu-performance-degradation" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation into .claude/skills/gke-ai-troubleshooting-tpu-performance-degradation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-performance-degradation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-performance-degradation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-performance-degradation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation .agents/skills/gke-ai-troubleshooting-tpu-performance-degradation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gke-ai-troubleshooting-tpu-performance-degradation" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation into .agents/skills/gke-ai-troubleshooting-tpu-performance-degradation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-performance-degradation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-performance-degradation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-performance-degradation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation .cursor/skills/gke-ai-troubleshooting-tpu-performance-degradation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gke-ai-troubleshooting-tpu-performance-degradation" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation into .cursor/skills/gke-ai-troubleshooting-tpu-performance-degradation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-performance-degradation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-performance-degradation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-performance-degradation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation .gemini/skills/gke-ai-troubleshooting-tpu-performance-degradation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-performance-degradation" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation into .gemini/skills/gke-ai-troubleshooting-tpu-performance-degradation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-performance-degradation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-performance-degradationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-performance-degradation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation .github/skills/gke-ai-troubleshooting-tpu-performance-degradation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-performance-degradation" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation into .github/skills/gke-ai-troubleshooting-tpu-performance-degradation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-performance-degradation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-tpu-performance-degradation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-tpu-performance-degradation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation .opencode/skills/gke-ai-troubleshooting-tpu-performance-degradation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-tpu-performance-degradation" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation into .opencode/skills/gke-ai-troubleshooting-tpu-performance-degradation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-tpu-performance-degradation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gke-ai-troubleshooting-tpu-performance-degradationDiagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (gcloud alpha mldiagnostics monitored-events /…
Gke AI Troubleshooting Tpu Performance Degradation is an agent skill from google/skills, published by the product's own GitHub organization. Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (gcloud alpha mldiagnostics monitored-events / hypercomputecluster.googleapis.com/v1alpha) and 1-minute Cloud Monitoring system metrics (kubernetes.io/node/accelerator/). Distinguishes hardware and network fabric throttling from workload resource bottlenecks (HBM capacity, host memory, or host CPU saturation). Use when TPU training throughput or duty cycle drops without…
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/path-a-workload-bottlenecks.md`, `references/path-b-infrastructure-throttling.md` and `references/path-c-no-detected-analyzer.md`).
It sits in DevOps & Cloud, covering Container orchestration. It works with Google Kubernetes Engine, Google Cloud and Kubernetes. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4b940dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gcloudkubectlFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.cloud.google.comcloud.google.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gke AI Troubleshooting Tpu Performance Degradation loads about 3.5k tokens when it runs, and up to ~5.1k if it reads all its reference files. Until then it costs about 226 tokens; SKILL.md has 1,179 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from google/skills at commit 4b940dd, republished under its Apache-2.0 licence (© google). 1,179 words, ~3,481 tokens.
.claude/skills/gke-ai-troubleshooting-tpu-performance-degradation/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Diagnose and mitigate Cloud TPU training throughput drops and step-time
regressions (15%+ drop in TPU duty cycle) on Google Kubernetes Engine (GKE) by
correlating ML Diagnostics Workload Monitoring MonitoredEvent analyzer
reports with 1-minute Cloud Monitoring system metrics and GKE node topology
labels.
gcloud with
alpha component for gcloud alpha mldiagnostics) and kubectl.gcloud billing projects describe {project_id}), authenticate
(gcloud auth login), set the target project (gcloud config set project {project_id}), and ensure container.googleapis.com,
monitoring.googleapis.com, and hypercomputecluster.googleapis.com are
enabled.jobset and job GKE
job types, and is compatible with GKE versions 1.36.0-gke.4681000 and later,
as stated in
Configure GKE for ML Diagnostics.
If a workload uses another framework (such as PyTorch) or another custom
resource type, gcloud alpha mldiagnostics won't list ML runs or monitored
events for it. On-demand profiling (Step 4 Path A) additionally requires the
cluster setup described in that document.roles/hypercomputecluster.editor), the role
listed in the "IAM permissions" section of
ML Diagnostics platform,
for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored
events, and on-demand profiler sessions)roles/monitoring.viewer) for the PromQL queries in Step
2roles/container.viewer) for the kubectl get nodes query in Step 3[High Risk] steps): Kubernetes Engine Cluster Admin
(roles/container.clusterAdmin)Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.
When you recommend a fix, link the doc section that describes it.
By default, Workload Monitoring treats a 15% drop in TPU duty cycle as a
performance degradation, raises a MonitoredEvent (type: PERFORMANCE_DEGRADATION), and runs its analyzers. Consult the
Workload Monitoring and analyzers
section in the official Cloud TPU documentation for the complete analyzer
definitions, detection criteria, and subsections, and route remediation by
category:
detectionState: "DETECTED") indicates a workload
memory or host CPU capacity bottleneck (see the "HBM Capacity Analyzer",
"Host Memory Utilization Analyzer", and "CPU Utilization Analyzer" sections
of
Workload monitoring with ML Diagnostics),
route to
Path A: Workload resource bottlenecks
(workload optimization and profiling).detectionState: "DETECTED") indicates an
interconnect, thermal/power throttling, memory bandwidth, or Top-of-Rack
(ToR) network fault on specific instances (see
Workload Monitoring and analyzers),
route to
Path B: Infrastructure or network fabric throttling
(map the culprit Compute Engine instance IDs to GKE nodes, handle node
repair through GKE, and escalate to Google Cloud Support).PERFORMANCE_DEGRADATION event fired, but no analyzer reports
detectionState: "DETECTED":PERFORMANCE_DEGRADATION event exists (as in the sample output in
List all monitored events,
where analyzerReports entries have detectionState: "NOT_DETECTED"),
route to
Path C: Event fired with no DETECTED analyzer.[Low Risk]Collect the target parameters. By default, query a 60-minute window `[T - 30m, T
around{issue_time}`:{project_id}: Google Cloud project ID{location}: Google Cloud region where the ML run and GKE cluster reside (for
example, us-central1){cluster_name}: GKE cluster name{workload_name} / {ml_run_id}: JobSet / workload name or ML Diagnostics
run ID{issue_time}: Timestamp when throughput degradation was observed (T,
ISO-8601 UTC){start_time}: T - 30m{end_time}: T + 30mmonitoredEvents [Low Risk]gcloud alpha mldiagnostics machine-learning-run list with a link to
List machine learning runs
in the ML Diagnostics CLI reference, and locate {ml_run_id} matching
{workload_name}.gcloud alpha mldiagnostics monitored-events list with a link to
Monitored-events commands,
or the hypercomputecluster.googleapis.com/v1alpha API with a link to
Access Workload Monitoring information through the API,
to check for PERFORMANCE_DEGRADATION events during [{start_time}, {end_time}].MonitoredEvent: Give the user gcloud alpha mldiagnostics monitored-events describe with a link to the same
Monitored-events commands
section, and inspect the analyzerReports array (analyzer,
detectionState, details, and recommendedActions).PERFORMANCE_DEGRADATION event fired, run Step 2 to corroborate with
the 1-minute system metrics; if none of its analyzerReports entries has
detectionState: "DETECTED", follow
Path C: Event fired with no DETECTED analyzer.PERFORMANCE_DEGRADATION event exists and duty cycle is steady in
Step 2, rule out TPU performance degradation by following
Path D: Healthy telemetry.[Low Risk]Consult the
System Metrics
section in the Workload Monitoring documentation for the 1-minute Cloud
Monitoring metrics exported for TPU and host devices, and run read-only PromQL
queries over [{start_time}, {end_time}] to corroborate the analyzer report:
# 1. Node TPU duty cycle (look for the drop on the affected nodes)
kubernetes_io:node_accelerator_duty_cycle{
monitored_resource="k8s_node",
project_id="{project_id}",
cluster_name="{cluster_name}"
}
# 2. HBM utilization ratio by node (around 0.90 is approaching the limit)
sum by (node_name) (
kubernetes_io:node_accelerator_memory_used{
monitored_resource="k8s_node",
project_id="{project_id}",
cluster_name="{cluster_name}"
}
)
/
sum by (node_name) (
kubernetes_io:node_accelerator_memory_total{
monitored_resource="k8s_node",
project_id="{project_id}",
cluster_name="{cluster_name}"
}
)
# 3. Host memory and CPU allocatable utilization
kubernetes_io:node_memory_allocatable_utilization{
monitored_resource="k8s_node",
project_id="{project_id}",
cluster_name="{cluster_name}"
}
kubernetes_io:node_cpu_allocatable_utilization{
monitored_resource="k8s_node",
project_id="{project_id}",
cluster_name="{cluster_name}"
}[Low Risk]When an infrastructure analyzer reports culprit numeric Compute Engine instance
IDs in details or recommendedActions, map those numeric instance IDs to GKE
Node names and physical topology blocks using this read-only kubectl query
inspecting container.googleapis.com/instance_id:
kubectl get nodes -l cloud.google.com/gke-tpu-accelerator \
-o jsonpath='{range .items[*]}{.metadata.name}{"\tinstance_id="}{.metadata.annotations.container\.googleapis\.com/instance_id}{"\tblock="}{.metadata.labels.cloud\.google\.com/gce-topology-block}{"\tsubblock="}{.metadata.labels.cloud\.google\.com/gce-topology-subblock}{"\thost="}{.metadata.labels.cloud\.google\.com/gce-topology-host}{"\n"}{end}'Load and follow only the reference file that matches the analyzer category from Step 1:
PERFORMANCE_DEGRADATION event fired with all analyzers
NOT_DETECTED): Read
Path C: Event fired with no DETECTED analyzer.PERFORMANCE_DEGRADATION events and steady telemetry): Read
Path D: Healthy telemetry.gcloud compute instance-groups managed commands, such as
delete, on a node pool's managed instance group. Handle nodes through GKE,
as described in
Path B: Infrastructure or network fabric throttling.© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation of google/skills.
Open the folder on GitHubat commit 4b940dd
Gke AI Troubleshooting Tpu Performance Degradation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gke AI Troubleshooting Tpu Performance Degradation this skillgoogle/skills | 21k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| Devopsnicepkg/auto-company | 195 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| Kcli Cluster Deploymentkarmab/kcli | 653 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| GCP Gkesickn33/agentic-awesome-skills | 47k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Apex Azure Cloud Migratejonathan-vella/apex | 217 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Dt Obs GCPDynatrace/dynatrace-for-ai | 163 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 |
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
sickn33/agentic-awesome-skills
Deploy and manage Google Kubernetes Engine clusters. An agent skill from sickn33/agentic-awesome-skills.
jonathan-vella/apex
WORKFLOW SKILL — Assess and migrate cross-cloud workloads to Azure: assessments and code conversion from AWS, GCP, Heroku, Kubernetes or Spring.
Dynatrace/dynatrace-for-ai
GCP cloud resources including Compute Engine, GKE, Cloud Run, Pub/Sub, VPC networking, DNS, IAM, Secret Manager, and monitoring.
metalbear-co/mirrord
Help users install and configure the mirrord Operator for team/enterprise environments.
google/skills
Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
Categories
Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (gcloud alpha mldiagnostics monitored-events /…. Gke AI Troubleshooting Tpu Performance Degradation is an agent skill from google/skills, published by the product's own GitHub organization.io/node/accelerator/).
Gke AI Troubleshooting Tpu Performance Degradation fits situations like: TPU training throughput; duty cycle drops without crashing pods; PERFORMANCEDEGRADATION monitored events fire; triaging slow multi-slice training steps.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-performance-degradation -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation in google/skills) into .claude/skills/gke-ai-troubleshooting-tpu-performance-degradation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-performance-degradation -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation in google/skills) into .agents/skills/gke-ai-troubleshooting-tpu-performance-degradation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-tpu-performance-degradation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-tpu-performance-degradation, .gemini/skills/gke-ai-troubleshooting-tpu-performance-degradation, .github/skills/gke-ai-troubleshooting-tpu-performance-degradation and .opencode/skills/gke-ai-troubleshooting-tpu-performance-degradation in your project.
Going by SKILL.md and its folder, Gke AI Troubleshooting Tpu Performance Degradation needs the command-line tools its instructions call (gcloud and kubectl).
SKILL.md names 2 domains. As links in the text: docs.cloud.google.com and cloud.google.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gke AI Troubleshooting Tpu Performance Degradation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Gke AI Troubleshooting Tpu Performance Degradation: Devops (nicepkg/auto-company, 195 stars), Kcli Cluster Deployment (karmab/kcli, 653 stars), GCP Gke (sickn33/agentic-awesome-skills, 47k stars) and Apex Azure Cloud Migrate (jonathan-vella/apex, 217 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.