Devops
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair.
$ npx skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-node-unresponsive-timeout --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout .claude/skills/gke-ai-troubleshooting-node-unresponsive-timeout && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gke-ai-troubleshooting-node-unresponsive-timeout" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout into .claude/skills/gke-ai-troubleshooting-node-unresponsive-timeout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-node-unresponsive-timeout", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeoutType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-node-unresponsive-timeout --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout .agents/skills/gke-ai-troubleshooting-node-unresponsive-timeout && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gke-ai-troubleshooting-node-unresponsive-timeout" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout into .agents/skills/gke-ai-troubleshooting-node-unresponsive-timeout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-node-unresponsive-timeout", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-node-unresponsive-timeout --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout .cursor/skills/gke-ai-troubleshooting-node-unresponsive-timeout && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gke-ai-troubleshooting-node-unresponsive-timeout" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout into .cursor/skills/gke-ai-troubleshooting-node-unresponsive-timeout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-node-unresponsive-timeout", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-node-unresponsive-timeout --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout .gemini/skills/gke-ai-troubleshooting-node-unresponsive-timeout && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-node-unresponsive-timeout" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout into .gemini/skills/gke-ai-troubleshooting-node-unresponsive-timeout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-node-unresponsive-timeout", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills gke-ai-troubleshooting-node-unresponsive-timeoutInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout .github/skills/gke-ai-troubleshooting-node-unresponsive-timeout && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-node-unresponsive-timeout" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout into .github/skills/gke-ai-troubleshooting-node-unresponsive-timeout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-node-unresponsive-timeout", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills gke-ai-troubleshooting-node-unresponsive-timeout --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout .opencode/skills/gke-ai-troubleshooting-node-unresponsive-timeout && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gke-ai-troubleshooting-node-unresponsive-timeout" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout into .opencode/skills/gke-ai-troubleshooting-node-unresponsive-timeout/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-ai-troubleshooting-node-unresponsive-timeout", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gke-ai-troubleshooting-node-unresponsive-timeoutDiagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair.
Gke AI Troubleshooting Node Unresponsive Timeout is an agent skill from google/skills, published by the product's own GitHub organization. Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair. Use when nodes stop heartbeating beyond the node auto-repair threshold and pods remain stuck in Terminating. Don't use for healthy nodes, pod-only application crashes, or routine GKE upgrades.
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud. It works with Google Kubernetes Engine, Google Cloud and Kubernetes. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 5120a76. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gcloudFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.cloud.google.comcloud.google.comkubernetes.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gke AI Troubleshooting Node Unresponsive Timeout loads about 3.2k tokens when it runs. Until then it costs about 109 tokens; SKILL.md has 1,269 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from google/skills at commit 5120a76, republished under its Apache-2.0 licence (© google). 1,269 words, ~3,183 tokens.
.claude/skills/gke-ai-troubleshooting-node-unresponsive-timeout/SKILL.md (or your agent's skills folder).NodeStatusUnknown)When the Compute Engine host of a TPU or GPU node has a fatal hardware error,
kernel panic, or non-maskable interrupt (NMI) lockup, the guest OS stops
responding. The kubelet can no longer send heartbeats, so the node Ready
condition becomes Unknown with Reason: NodeStatusUnknown (Kubelet stopped posting node status.). If node auto-repair is disabled on the node pool, GKE
doesn't repair the node. The node can stay NotReady, and pods on it can stay
in Terminating, which blocks multi-host JobSet workloads from recovering.
gcloud) and
kubectl.gcloud billing projects describe {project_id}), authenticate
(gcloud auth login), set the target project (gcloud config set project {project_id}), and ensure container.googleapis.com,
compute.googleapis.com, logging.googleapis.com, and
monitoring.googleapis.com are enabled.roles/container.viewer)roles/compute.viewer)roles/logging.viewer)roles/monitoring.viewer)[High Risk] steps): Kubernetes Engine Cluster Admin
(roles/container.clusterAdmin)Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.
When you recommend a fix, link the doc section that describes it.
[Low Risk]Collect the target parameters. By default, query a 60-minute window `[T - 30m, T
around{issue_time}`:{project_id}: Google Cloud project ID{cluster_name}: GKE cluster name{location}: Cluster region or zone{nodepool_name}: Target TPU or GPU node pool name{node_name}: Unresponsive GKE node name (and its Compute Engine {zone}){issue_time}: Incident timestamp in RFC3339 UTC{start_time}: {issue_time} - 30m{end_time}: {issue_time} + 30mNodeStatusUnknown heartbeat timeout [Low Risk]Ready condition is Unknown with Reason: NodeStatusUnknown
(Kubelet stopped posting node status.), follow the instructions in the
section
Check the node's status and conditions.k8s_node and k8s_cluster
logs across [{start_time}, {end_time}] to confirm when the control plane
lost heartbeat contact with {node_name}:(resource.type="k8s_node" OR resource.type="k8s_cluster")
resource.labels.cluster_name="{cluster_name}"
("{node_name}" AND ("NodeNotReady" OR "NodeStatusUnknown" OR "Kubelet stopped posting node status"))
timestamp >= "{start_time}" AND timestamp <= "{end_time}"Unknown state using the GKE system metric
kubernetes.io/node/status_condition (kubernetes_io:node_status_condition,
GKE 1.32.1-gke.1357001+) documented in
Monitor health metrics for TPU nodes and node pools,
filtered by condition="Ready" and status="Unknown":kubernetes_io:node_status_condition{
monitored_resource="k8s_node",
cluster_name="{cluster_name}",
node_name="{node_name}",
condition="Ready",
status="Unknown"
}Ready condition is True and NodeStatusUnknown is absent,
rule out an unresponsive node timeout and pivot to workload-level
troubleshooting (for example,
Troubleshoot OOM events)
rather than repairing or draining the node.Ready is Unknown (NodeStatusUnknown), proceed to Step 2.[Low Risk]The guest OS can't send logs after a fatal kernel freeze, so check the serial port output of the node's VM:
serial-port-logging-enable is set to true), GKE system logs in Cloud
Logging include the node's serial port output. See the section
System logs.
Use Cloud Logging when the VM is stopped or has already been replaced by
auto-repair, or when you need more than the most recent output.gcloud compute instances get-serial-port-output with --port=1) for {node_name} in {zone}. This
method returns only the most recent 1 MB of output per port.Consult
Troubleshoot Linux VM boot issues due to kernel panic
to identify documented kernel panic and hardware crash patterns (such as Fatal Machine check, hung_task: blocked tasks, or NMI: Not continuing) in the
serial port output.
[Low Risk]Find out why GKE hasn't repaired the unresponsive node:
autoRepair is enabled on {nodepool_name}.[Low Risk]To list the system events for {node_name} around {issue_time}, follow the
instructions in the section
Querying Cloud Audit Logs.
Compare the method field with the table in the "Reviewing Cloud Audit Logs"
section of the same document, and look for:
compute.instances.hostError: a hardware or software issue on the physical
host caused the VM to crash.compute.instances.preempted: Compute Engine preempted a Spot VM or
preemptible VM. For preempted nodes, also see
Confirm node preemption.compute.instances.automaticRestart: Compute Engine restarted the VM after a
hostError or terminateOnHostMaintenance event.compute.instances.guestTerminate: the VM's operating system initiated the
shutdown.[High Risk]Guardrails:
- Never force-delete stuck
Terminatingpods on an unresponsive node. Force deletion doesn't wait for the kubelet to confirm that the pod has stopped, so a replacement pod can start while the old one is still running. See Force Delete StatefulSet Pods.- Never delete GKE-managed Compute Engine VM instances directly (
gcloud compute instances delete). Instead, check how long the node has reportedNodeStatusUnknown(Step 1), check the serial console output (Step 2), and rely on node auto-repair, as described in the following steps.
autoRepair is disabled on {nodepool_name}, link the user to
Enable auto-repair for an existing Standard node pool
and
Configure auto repair for TPU slice nodes
so they can apply the change themselves.NotReady or no status for
the documented time threshold by draining and re-creating it, and link the
"Repair criteria" and "Node repair process" sections of
Auto-repair nodes.
GKE waits one hour for the drain to complete. If the drain doesn't
complete, GKE shuts the node down and creates a new node. Tell the user to
expect this one-hour drain wait before GKE re-creates the node.Ready status, then confirm
that the JobSet pods are running again.© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout of google/skills.
Open the folder on GitHubat commit 5120a76
Gke AI Troubleshooting Node Unresponsive Timeout next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gke AI Troubleshooting Node Unresponsive Timeout this skillgoogle/skills | 21k | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| Devopsnicepkg/auto-company | 194 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| Kcli Cluster Deploymentkarmab/kcli | 653 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| GCP Gkesickn33/agentic-awesome-skills | 47k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Apex Azure Cloud Migratejonathan-vella/apex | 217 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Dt Obs GCPDynatrace/dynatrace-for-ai | 162 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 |
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
sickn33/agentic-awesome-skills
Deploy and manage Google Kubernetes Engine clusters. An agent skill from sickn33/agentic-awesome-skills.
jonathan-vella/apex
WORKFLOW SKILL — Assess and migrate cross-cloud workloads to Azure: assessments and code conversion from AWS, GCP, Heroku, Kubernetes or Spring.
Dynatrace/dynatrace-for-ai
GCP cloud resources including Compute Engine, GKE, Cloud Run, Pub/Sub, VPC networking, DNS, IAM, Secret Manager, and monitoring.
maslennikov-ig/claude-code-orchestrator-kit
Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…
google/skills
Query Cloud Trace spans, filter by latency thresholds or error status, correlate distributed traces with Cloud Logging, and diagnose latency bottlenecks across Google Cloud services.
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
Categories
Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair. Gke AI Troubleshooting Node Unresponsive Timeout is an agent skill from google/skills, published by the product's own GitHub organization. Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair.
Gke AI Troubleshooting Node Unresponsive Timeout fits situations like: nodes stop heartbeating beyond the node auto-repair threshold and pods remain stuck in Terminating; pod-only application crashes; routine GKE upgrades.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout -a claude-code`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout in google/skills) into .claude/skills/gke-ai-troubleshooting-node-unresponsive-timeout in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout -a codex`. Or copy the skill folder (skills/cloud/gke-ai-troubleshooting-node-unresponsive-timeout in google/skills) into .agents/skills/gke-ai-troubleshooting-node-unresponsive-timeout in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-ai-troubleshooting-node-unresponsive-timeout, .gemini/skills/gke-ai-troubleshooting-node-unresponsive-timeout, .github/skills/gke-ai-troubleshooting-node-unresponsive-timeout and .opencode/skills/gke-ai-troubleshooting-node-unresponsive-timeout in your project.
Going by SKILL.md and its folder, Gke AI Troubleshooting Node Unresponsive Timeout needs the command-line tools its instructions call (gcloud).
SKILL.md names 3 domains. As links in the text: docs.cloud.google.com, cloud.google.com and kubernetes.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gke AI Troubleshooting Node Unresponsive Timeout is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Gke AI Troubleshooting Node Unresponsive Timeout: Devops (nicepkg/auto-company, 194 stars), Kcli Cluster Deployment (karmab/kcli, 653 stars), GCP Gke (sickn33/agentic-awesome-skills, 47k stars) and Apex Azure Cloud Migrate (jonathan-vella/apex, 217 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 21,069 GitHub stars. The repository holds 147 skills in this directory. The repository was last updated on October 9, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.