Devops
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
Diagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations.
$ npx skills add google/skills --skill gke-node-notready -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills gke-node-notready --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/gke-node-notready .claude/skills/gke-node-notready && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gke-node-notready" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-node-notready into .claude/skills/gke-node-notready/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-node-notready", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/gke-node-notreadyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill gke-node-notready -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills gke-node-notready --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/gke-node-notready .agents/skills/gke-node-notready && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gke-node-notready" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-node-notready into .agents/skills/gke-node-notready/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-node-notready", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-node-notready -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills gke-node-notready --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/gke-node-notready .cursor/skills/gke-node-notready && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gke-node-notready" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-node-notready into .cursor/skills/gke-node-notready/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-node-notready", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/gke-node-notready--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill gke-node-notready -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills gke-node-notready --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/gke-node-notready .gemini/skills/gke-node-notready && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gke-node-notready" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-node-notready into .gemini/skills/gke-node-notready/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-node-notready", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills gke-node-notreadyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill gke-node-notready -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/gke-node-notready .github/skills/gke-node-notready && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gke-node-notready" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-node-notready into .github/skills/gke-node-notready/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-node-notready", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill gke-node-notready -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills gke-node-notready --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/gke-node-notready .opencode/skills/gke-node-notready && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gke-node-notready" agent skill from https://github.com/google/skills/tree/main/skills/cloud/gke-node-notready into .opencode/skills/gke-node-notready/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gke-node-notready", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gke-node-notreadyDiagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations.
Gke Node Notready is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations. Use when nodes show NotReady, when the kubelet stops posting node status, or when workloads are evicted or stuck Pending due to node health. Don't use for pod-level application failures (use gke-workload-troubleshooting), autoscaler scale-up/scale-down decisions (use gke-cluster-autoscaler), or non-GKE compute.
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud. It works with Google Kubernetes Engine, Kubernetes and Google Cloud. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8a1ac05. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlgcloudFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
console.cloud.google.comAlso links to:
cloud.google.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gke Node Notready loads about 2.9k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,150 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from google/skills at commit 8a1ac05, republished under its Apache-2.0 licence (© google). 1,150 words, ~2,942 tokens.
.claude/skills/gke-node-notready/SKILL.md (or your agent's skills folder).Use this skill to systematically diagnose why one or more GKE nodes report a
NotReady (or Ready: Unknown) status and to propose safe remediations. A
NotReady status means the node's kubelet is not reporting to the control plane
correctly, so Kubernetes stops scheduling new Pods on the node, which can reduce
application capacity and cause downtime.
This skill operates non-interactively and enforces a read-only diagnostics
boundary: gather evidence first, then propose a fix (a kubectl/gcloud
command or a GitOps manifest change) for a human to apply. Never mutate the
cluster, drain, delete, or recreate nodes automatically.
[!IMPORTANT] First rule out an expected
NotReady: a node that is newly provisioning, upgrading, being repaired, cordoned, or scaling down will transiently reportNotReady. Only treat it as a fault if it persists beyond the expected window.
project_id, cluster_name,
cluster_location, and node_name non-interactively from the user prompt,
active SETTINGS.md, or environment defaults (kubectl config current-context,
gcloud config get-value project).gcloud container clusters get-credentials {cluster_name} --location {cluster_location} --project {project_id}.
If the cluster is unreachable or commands fail (sandbox/dry-run/offline),
present the exact diagnostic commands for a human to run and continue the
analysis from the reported symptoms.{issue_time} (explicit, relative, or now) and
center a 1-hour window around it (start = issue_time - 30m,
end = issue_time + 30m) for all log/metric queries.# List nodes and spot NotReady status, node IPs, and container-runtime version.
kubectl get nodes -o wide
# Inspect the affected node's Conditions and Events (the primary clues).
kubectl describe node "{node_name}"Equivalent via Cloud Logging (preferred when kubectl access is limited or for
historical events). Open it as a Logs Explorer deep link — URL-encode the
query and append the project and Step 0 time window:
https://console.cloud.google.com/logs/query;query={URL_ENCODED_QUERY};timeRange={start}%2F{end}?project={project_id}
(encode / as %2F, or use ;duration=PT1H for a rolling hour):
resource.type="k8s_node"
log_id("events")
resource.labels.node_name="{node_name}"
resource.labels.cluster_name="{cluster_name}"
resource.labels.location="{cluster_location}"Interpret the Conditions table:
Ready: False / Ready: Unknown with reason KubeletNotReady /
NodeStatusUnknown ("Kubelet stopped posting node status") → kubelet or
runtime problem; continue to Step 2.MemoryPressure: True, DiskPressure: True, PIDPressure: True → resource
exhaustion; go to Step 4b.NetworkUnavailable: True → networking/CNI problem; go to Step 4d.Open these kubelet logs as a Logs Explorer deep link using the same
logs/query;query={URL_ENCODED_QUERY};timeRange=...?project=... pattern as Step 1.
resource.type="k8s_node"
resource.labels.node_name="{node_name}"
resource.labels.cluster_name="{cluster_name}"
resource.labels.location="{cluster_location}"
log_id("kubelet")
severity>=WARNINGAlso review the node's serial-console logs (log_id("serialconsole.googleapis.com/serial_port_1_output")
or the resource.type="gce_instance" serial logs) for kernel TaskHung,
OOM-killer, or disk I/O errors that correlate with the kubelet failures.
| Kubelet / event signature | Likely root cause | Go to |
|---|---|---|
runtime is down, Container runtime not ready, errors on /run/containerd/containerd.sock (connection refused / DeadlineExceeded) | Container runtime (containerd) down or unresponsive | Step 4a |
Got sys oom event from cadvisor / kernel OOM-killer in serial logs | System (node-level) OOM killed critical processes | Step 4b |
PLEG is not healthy | PLEG stalled, usually node overload (CPU/disk) | Step 4c |
TaskHung for containerd/kubelet, high disk latency | Disk throttling / I/O starvation | Step 4b |
failed to ensure lease, leases.coordination.k8s.io ... namespace kube-node-lease ... terminating | kube-node-lease termination → NotReady flapping | Step 4f |
| Kubelet cannot reach API server, TLS/dial timeouts | Kubelet ↔ control-plane connectivity | Step 4d |
NetworkPluginNotReady, cni plugin not initialized, NetworkUnavailable | CNI plugin failure | Step 4d |
| Node-critical DaemonSet Pods (CNI, kube-proxy, metadata) blocked from admission | Admission webhook interference | Step 4e |
Only generic NodeNotReady, no other signature | Cause unclear — widen to Step 4d, then escalate | Escalation |
containerd) downConfirm the kubelet cannot talk to containerd (socket errors above). Check for
containerd restarts/crashes in serial logs. Remediation (propose, don't run):
recreate/repair the node (kubectl drain then let the node pool recreate it, or
gcloud container clusters upgrade/node auto-repair); if it recurs across nodes,
suspect a node image or custom DaemonSet interfering with containerd.
# Node allocatable vs. usage.
kubectl describe node "{node_name}" | sed -n '/Allocated resources/,/Events/p'Cloud Monitoring metrics to inspect (read-only): kubernetes.io/node/memory/used_bytes,
kubernetes.io/node/cpu/core_usage_time, kubernetes.io/node/ephemeral_storage/used_bytes.
requests/limits,
reduce over-commit, or use larger machine types. Distinguish system OOM
(node-wide, kills kubelet/runtime) from cgroup OOM (single container).PLEG is not healthy almost always means the node is overloaded (CPU saturation,
disk latency, or too many Pods/containers per node) so the runtime can't relist
in time. Correlate with 4b metrics. Remediation: reduce node density, add
CPU/disk headroom, or spread workloads.
# Are node-critical networking Pods healthy on this node?
kubectl get pods -n kube-system -o wide --field-selector spec.nodeName={node_name}NetworkPluginNotReady): the CNI DaemonSet
(netd/calico/dataplane) is not running on the node → inspect those Pods'
logs/events.A misconfigured/failing validating or mutating webhook with a broad scope can block node-critical system Pods from being admitted, keeping the node NotReady.
kubectl get validatingwebhookconfigurations,mutatingwebhookconfigurationsLook for webhooks that intercept kube-system / node-critical objects with
failurePolicy: Fail. Remediation (propose): scope the webhook out of
kube-system/node-critical namespaces or set an appropriate namespaceSelector.
kube-node-lease termination flappingIf the node flaps NotReady with leases.coordination.k8s.io ... namespace kube-node-lease ... is being terminated, the kube-node-lease namespace was
deleted/terminating. Remediation (propose): do not delete the
kube-node-lease namespace; if terminating, identify the finalizer/actor holding
it and restore the namespace.
Escalate when either:
_Default bucket defaults to 30 days, so
incidents older than that are permanently deleted); orIn those cases, do all three:
_Default retention window and are permanently deleted")._Required bucket, default
400-day retention).This skill is derived from public Google Cloud documentation:
kube-node-lease / CNI / admission-webhook signatures and their remediations.resource.type="k8s_node", log_id("kubelet")) and log-bucket
retention (_Default 30 days, _Required 400 days).logs/query;query=... deep-link
format used above).© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/cloud/gke-node-notready of google/skills.
Open the folder on GitHubat commit 8a1ac05
Gke Node Notready next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gke Node Notready this skillgoogle/skills | 21k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Devopsnicepkg/auto-company | 192 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| Kcli Cluster Deploymentkarmab/kcli | 653 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| GCP Gkesickn33/agentic-awesome-skills | 47k | 2 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Apex Azure Cloud Migratejonathan-vella/apex | 217 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Dt Obs GCPDynatrace/dynatrace-for-ai | 161 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 |
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
sickn33/agentic-awesome-skills
Deploy and manage Google Kubernetes Engine clusters. An agent skill from sickn33/agentic-awesome-skills.
jonathan-vella/apex
WORKFLOW SKILL — Assess and migrate cross-cloud workloads to Azure: assessments and code conversion from AWS, GCP, Heroku, Kubernetes or Spring.
Dynatrace/dynatrace-for-ai
GCP cloud resources including Compute Engine, GKE, Cloud Run, Pub/Sub, VPC networking, DNS, IAM, Secret Manager, and monitoring.
maslennikov-ig/claude-code-orchestrator-kit
Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
google/skills
Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans.
Categories
Diagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations. Gke Node Notready is an agent skill from google/skills, published by the product's own GitHub organization. Diagnoses GKE nodes reporting NotReady or Unknown status by inspecting node conditions, events, kubelet/containerd logs, and node metrics, then proposing safe remediations.
Gke Node Notready fits situations like: nodes show NotReady; the kubelet stops posting node status; workloads are evicted; stuck Pending due to node health.
Run `npx skills add google/skills --skill gke-node-notready -a claude-code`. Or copy the skill folder (skills/cloud/gke-node-notready in google/skills) into .claude/skills/gke-node-notready in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill gke-node-notready -a codex`. Or copy the skill folder (skills/cloud/gke-node-notready in google/skills) into .agents/skills/gke-node-notready in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill gke-node-notready -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gke-node-notready, .gemini/skills/gke-node-notready, .github/skills/gke-node-notready and .opencode/skills/gke-node-notready in your project.
Going by SKILL.md and its folder, Gke Node Notready needs the command-line tools its instructions call (kubectl and gcloud).
SKILL.md names 2 domains. In commands or code: console.cloud.google.com; the agent is likely to contact it when it follows the instructions. As links in the text: cloud.google.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gke Node Notready is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Gke Node Notready: Devops (nicepkg/auto-company, 192 stars), Kcli Cluster Deployment (karmab/kcli, 653 stars), GCP Gke (sickn33/agentic-awesome-skills, 47k stars) and Apex Azure Cloud Migrate (jonathan-vella/apex, 217 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 20,994 GitHub stars. The repository holds 145 skills in this directory. The repository was last updated on October 6, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.