Official agent skill

Gke Node Notready

by google in google/skills

Diagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Gke Node Notready

skills CLI
$ npx skills add google/skills --skill gke-node-notready -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google/skills gke-node-notready --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-node-notready .claude/skills/gke-node-notready && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gke-node-notready
GitHub stars
21k
Token cost
~2.9k tokens
SKILL.md length
1,150 words
Files
1
Skills in repo
145
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations.

  • Works in 6 steps: Context discovery & time window → Identify NotReady nodes and gather… → Scan kubelet logs for error signatures → …
  • Nodes show NotReady
  • SKILL.md covers 🔍 Diagnostic Workflow and References
  • Calls kubectl and gcloud; reaches console.cloud.google.com

What it does

Gke Node Notready is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations. Use when nodes show NotReady, when the kubelet stops posting node status, or when workloads are evicted or stuck Pending due to node health. Don't use for pod-level application failures (use gke-workload-troubleshooting), autoscaler scale-up/scale-down decisions (use gke-cluster-autoscaler), or non-GKE compute.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud. It works with Google Kubernetes Engine, Kubernetes and Google Cloud. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.

When your agent uses it

  • Nodes show NotReady
  • The kubelet stops posting node status
  • Workloads are evicted
  • Stuck Pending due to node health

Example prompts

  • “Use the gke-node-notready skill to diagnose GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd…”
  • “/gke-node-notready”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Context discovery & time window
  2. Identify NotReady nodes and gather initial status
  3. Scan kubelet logs for error signatures
  4. Map the signature to a root cause (decision table)
  5. Branch investigations
  6. Remediation boundary & escalation

What it can do on your machine

Read from SKILL.md and the folder at commit 8a1ac05. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • gcloud

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • console.cloud.google.com

    Also links to:

    • cloud.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gke Node Notready loads about 2.9k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,150 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from google/skills at commit 8a1ac05, republished under its Apache-2.0 licence (© google). 1,150 words, ~2,942 tokens.

Download SKILL.mdSave it as .claude/skills/gke-node-notready/SKILL.md (or your agent's skills folder).
name
gke-node-notready
description
Diagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations. Use when nodes show NotReady, when the kubelet stops posting node status, or when workloads are evicted or stuck Pending due to node health. Don't use for pod-level application failures (use gke-workload-troubleshooting), autoscaler scale-up/scale-down decisions (use gke-cluster-autoscaler), or non-GKE compute.
metadata.version
1.0.0
metadata.category
Containers

GKE Node NotReady Troubleshooting Skill

Use this skill to systematically diagnose why one or more GKE nodes report a NotReady (or Ready: Unknown) status and to propose safe remediations. A NotReady status means the node's kubelet is not reporting to the control plane correctly, so Kubernetes stops scheduling new Pods on the node, which can reduce application capacity and cause downtime.

This skill operates non-interactively and enforces a read-only diagnostics boundary: gather evidence first, then propose a fix (a kubectl/gcloud command or a GitOps manifest change) for a human to apply. Never mutate the cluster, drain, delete, or recreate nodes automatically.

[!IMPORTANT] First rule out an expected NotReady: a node that is newly provisioning, upgrading, being repaired, cordoned, or scaling down will transiently report NotReady. Only treat it as a fault if it persists beyond the expected window.

🔍 Diagnostic Workflow

Step 0: Context discovery & time window
  1. Parameter extraction — obtain project_id, cluster_name, cluster_location, and node_name non-interactively from the user prompt, active SETTINGS.md, or environment defaults (kubectl config current-context, gcloud config get-value project).
  2. Credentials & fallback — attempt gcloud container clusters get-credentials {cluster_name} --location {cluster_location} --project {project_id}. If the cluster is unreachable or commands fail (sandbox/dry-run/offline), present the exact diagnostic commands for a human to run and continue the analysis from the reported symptoms.
  3. Time window — determine {issue_time} (explicit, relative, or now) and center a 1-hour window around it (start = issue_time - 30m, end = issue_time + 30m) for all log/metric queries.

Step 1: Identify NotReady nodes and gather initial status
bash
# List nodes and spot NotReady status, node IPs, and container-runtime version.
kubectl get nodes -o wide

# Inspect the affected node's Conditions and Events (the primary clues).
kubectl describe node "{node_name}"

Equivalent via Cloud Logging (preferred when kubectl access is limited or for historical events). Open it as a Logs Explorer deep link — URL-encode the query and append the project and Step 0 time window: https://console.cloud.google.com/logs/query;query={URL_ENCODED_QUERY};timeRange={start}%2F{end}?project={project_id} (encode / as %2F, or use ;duration=PT1H for a rolling hour):

resource.type="k8s_node"
log_id("events")
resource.labels.node_name="{node_name}"
resource.labels.cluster_name="{cluster_name}"
resource.labels.location="{cluster_location}"

Interpret the Conditions table:

  • Ready: False / Ready: Unknown with reason KubeletNotReady / NodeStatusUnknown ("Kubelet stopped posting node status") → kubelet or runtime problem; continue to Step 2.
  • MemoryPressure: True, DiskPressure: True, PIDPressure: True → resource exhaustion; go to Step 4b.
  • NetworkUnavailable: True → networking/CNI problem; go to Step 4d.

Step 2: Scan kubelet logs for error signatures

Open these kubelet logs as a Logs Explorer deep link using the same logs/query;query={URL_ENCODED_QUERY};timeRange=...?project=... pattern as Step 1.

resource.type="k8s_node"
resource.labels.node_name="{node_name}"
resource.labels.cluster_name="{cluster_name}"
resource.labels.location="{cluster_location}"
log_id("kubelet")
severity>=WARNING

Also review the node's serial-console logs (log_id("serialconsole.googleapis.com/serial_port_1_output") or the resource.type="gce_instance" serial logs) for kernel TaskHung, OOM-killer, or disk I/O errors that correlate with the kubelet failures.


Step 3: Map the signature to a root cause (decision table)
Kubelet / event signatureLikely root causeGo to
runtime is down, Container runtime not ready, errors on /run/containerd/containerd.sock (connection refused / DeadlineExceeded)Container runtime (containerd) down or unresponsiveStep 4a
Got sys oom event from cadvisor / kernel OOM-killer in serial logsSystem (node-level) OOM killed critical processesStep 4b
PLEG is not healthyPLEG stalled, usually node overload (CPU/disk)Step 4c
TaskHung for containerd/kubelet, high disk latencyDisk throttling / I/O starvationStep 4b
failed to ensure lease, leases.coordination.k8s.io ... namespace kube-node-lease ... terminatingkube-node-lease termination → NotReady flappingStep 4f
Kubelet cannot reach API server, TLS/dial timeoutsKubelet ↔ control-plane connectivityStep 4d
NetworkPluginNotReady, cni plugin not initialized, NetworkUnavailableCNI plugin failureStep 4d
Node-critical DaemonSet Pods (CNI, kube-proxy, metadata) blocked from admissionAdmission webhook interferenceStep 4e
Only generic NodeNotReady, no other signatureCause unclear — widen to Step 4d, then escalateEscalation

Step 4: Branch investigations
4a. Container runtime (containerd) down

Confirm the kubelet cannot talk to containerd (socket errors above). Check for containerd restarts/crashes in serial logs. Remediation (propose, don't run): recreate/repair the node (kubectl drain then let the node pool recreate it, or gcloud container clusters upgrade/node auto-repair); if it recurs across nodes, suspect a node image or custom DaemonSet interfering with containerd.

4b. Resource pressure & OOM
bash
# Node allocatable vs. usage.
kubectl describe node "{node_name}" | sed -n '/Allocated resources/,/Events/p'

Cloud Monitoring metrics to inspect (read-only): kubernetes.io/node/memory/used_bytes, kubernetes.io/node/cpu/core_usage_time, kubernetes.io/node/ephemeral_storage/used_bytes.

  • DiskPressure / disk throttling: full boot disk or slow PD → increase disk size / use a faster PD type; reduce image/log churn.
  • System OOM: node memory exhausted → set/raise Pod memory requests/limits, reduce over-commit, or use larger machine types. Distinguish system OOM (node-wide, kills kubelet/runtime) from cgroup OOM (single container).
  • PIDPressure: too many processes → cap Pod PIDs / reduce workload density.
Show full SKILL.md (479 more words)Show less
4c. PLEG is not healthy

PLEG is not healthy almost always means the node is overloaded (CPU saturation, disk latency, or too many Pods/containers per node) so the runtime can't relist in time. Correlate with 4b metrics. Remediation: reduce node density, add CPU/disk headroom, or spread workloads.

4d. Networking
bash
# Are node-critical networking Pods healthy on this node?
kubectl get pods -n kube-system -o wide --field-selector spec.nodeName={node_name}
  • Kubelet ↔ control-plane: dial/TLS timeouts to the API server → check firewall rules, Private Google Access, authorized networks, and route/NAT changes.
  • CNI failure (NetworkPluginNotReady): the CNI DaemonSet (netd/calico/dataplane) is not running on the node → inspect those Pods' logs/events.
4e. Admission webhook interference

A misconfigured/failing validating or mutating webhook with a broad scope can block node-critical system Pods from being admitted, keeping the node NotReady.

bash
kubectl get validatingwebhookconfigurations,mutatingwebhookconfigurations

Look for webhooks that intercept kube-system / node-critical objects with failurePolicy: Fail. Remediation (propose): scope the webhook out of kube-system/node-critical namespaces or set an appropriate namespaceSelector.

4f. kube-node-lease termination flapping

If the node flaps NotReady with leases.coordination.k8s.io ... namespace kube-node-lease ... is being terminated, the kube-node-lease namespace was deleted/terminating. Remediation (propose): do not delete the kube-node-lease namespace; if terminating, identify the finalizer/actor holding it and restore the namespace.


Step 5: Remediation boundary & escalation
  • Present the root cause + evidence (the exact conditions, events, log lines, or metrics observed). Provide Cloud Logging deep links (and Cloud Monitoring links for the Step 4b metrics) to the supporting entries — using the deep-link pattern from Steps 1-2 — so a human can open the evidence directly.
  • Propose the fix as a command or GitOps manifest change for a human to apply — never apply, drain, or recreate nodes automatically. When to escalate (do this instead of proposing more self-service diagnostics):

Escalate when either:

  • the relevant logs are unavailable — excluded by a logging filter, or older than the log bucket's retention (the _Default bucket defaults to 30 days, so incidents older than that are permanently deleted); or
  • the kubelet/event signature is not in the Step 3 table and the root cause remains undetermined after the branch investigations.

In those cases, do all three:

  1. State the limitation plainly (for example, "kubelet logs for that date are past the 30-day _Default retention window and are permanently deleted").
  2. Summarize the findings you did gather (node conditions, events, metrics, and any Admin Activity audit logs still in the _Required bucket, default 400-day retention).
  3. Route to GKE support / engineering escalation with those findings. Do not keep proposing further self-service investigation, and do not fabricate a diagnosis when the evidence is missing.

References

This skill is derived from public Google Cloud documentation:

© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/cloud/gke-node-notready of google/skills.

Open the folder on GitHubat commit 8a1ac05

Compare with similar skills

Gke Node Notready next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gke Node Notready compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gke Node Notready this skillgoogle/skills21k—~2.9kAutomated safety check: PassApache-2.0
Devopsnicepkg/auto-company1922 repos~814Automated safety check: PassMIT
Kcli Cluster Deploymentkarmab/kcli653—~1.5kAutomated safety check: PassApache-2.0
GCP Gkesickn33/agentic-awesome-skills47k2 repos~2.5kAutomated safety check: PassMIT
Apex Azure Cloud Migratejonathan-vella/apex217—~1.2kAutomated safety check: PassMIT
Dt Obs GCPDynatrace/dynatrace-for-ai161—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    192 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • Guides deployment and management of Kubernetes clusters with kcli.

    653 GitHub stars~1.5k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • GCP Gke

    sickn33/agentic-awesome-skills

    Deploy and manage Google Kubernetes Engine clusters. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~2.5k tokens
    DevOps & CloudAuto-check passed
  • Apex Azure Cloud Migrate

    jonathan-vella/apex

    WORKFLOW SKILL — Assess and migrate cross-cloud workloads to Azure: assessments and code conversion from AWS, GCP, Heroku, Kubernetes or Spring.

    217 GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Dt Obs GCP

    Dynatrace/dynatrace-for-ai

    GCP cloud resources including Compute Engine, GKE, Cloud Run, Pub/Sub, VPC networking, DNS, IAM, Secret Manager, and monitoring.

    161 GitHub stars~2.5k tokensUpdated 6 days ago
    DevOps & CloudAuto-check passed
  • Senior DevOps Toolkit

    maslennikov-ig/claude-code-orchestrator-kit

    Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…

    259 GitHub starsUsed in 6 repos~1.1k tokens
    DevOps & CloudAuto-check: notes

More from google/skills

All 145 skills in this repo
  • Official

    Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.

    21k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

    21k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Official

    Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.

    21k GitHub stars~584 tokensUpdated today
    Auto-check passed
  • Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.

    21k GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Official

    Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans.

    21k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Categories

Questions about Gke Node Notready

What does Gke Node Notready do?

Diagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations. Gke Node Notready is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations.

When should I use Gke Node Notready?

Gke Node Notready fits situations like: nodes show NotReady; the kubelet stops posting node status; workloads are evicted; stuck Pending due to node health.

How do I install Gke Node Notready in Claude Code?

Run `npx skills add google/skills --skill gke-node-notready -a claude-code`. Or copy the skill folder (skills/cloud/gke-node-notready in google/skills) into .claude/skills/gke-node-notready in your project. Claude Code loads it when a task matches its description.

How do I install Gke Node Notready in Codex?

Run `npx skills add google/skills --skill gke-node-notready -a codex`. Or copy the skill folder (skills/cloud/gke-node-notready in google/skills) into .agents/skills/gke-node-notready in your project. Codex loads it when a task matches its description.

Can I use Gke Node Notready in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-node-notready -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-node-notready, .gemini/skills/gke-node-notready, .github/skills/gke-node-notready and .opencode/skills/gke-node-notready in your project.

What does Gke Node Notready need to run?

Going by SKILL.md and its folder, Gke Node Notready needs the command-line tools its instructions call (kubectl and gcloud).

Does Gke Node Notready access the network?

SKILL.md names 2 domains. In commands or code: console.cloud.google.com; the agent is likely to contact it when it follows the instructions. As links in the text: cloud.google.com. This is read from the text; nothing was executed.

Is Gke Node Notready safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gke Node Notready use?

Gke Node Notready is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gke Node Notready use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gke Node Notready?

Skills that share tags, products or a category with Gke Node Notready: Devops (nicepkg/auto-company, 192 stars), Kcli Cluster Deployment (karmab/kcli, 653 stars), GCP Gke (sickn33/agentic-awesome-skills, 47k stars) and Apex Azure Cloud Migrate (jonathan-vella/apex, 217 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gke Node Notready?

google (a GitHub organization, an official publisher) maintains it in google/skills, which has 20,994 GitHub stars. The repository holds 145 skills in this directory. The repository was last updated on October 6, 2026.

Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.