Agent skill

K8s Debug

by akin-ozer in akin-ozer/cc-devops-skills

Diagnose and fix Kubernetes pods, CrashLoopBackOff, Pending, DNS, networking, storage, and rollout failures with kubectl.

Apache-2.0Auto-check passedDevOps & Cloud

Install K8s Debug

skills CLI
$ npx skills add akin-ozer/cc-devops-skills --skill k8s-debug -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install akin-ozer/cc-devops-skills k8s-debug --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/akin-ozer/cc-devops-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/devops-skills-plugin/skills/k8s-debug .claude/skills/k8s-debug && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
k8s-debug
GitHub stars
319
Token cost
~2.9k tokens
SKILL.md length
895 words
Files
8 (incl. scripts, references)
Skills in repo
30
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnose and fix Kubernetes pods, CrashLoopBackOff, Pending, DNS, networking, storage, and rollout failures with kubectl.

  • Works in 6 steps: Preflight and Scope → Identify the Problem Layer → Gather Diagnostics with the Right Script → …
  • Tasks that involve Container orchestration
  • SKILL.md covers Overview, Trigger Phrases, Prerequisites and When to Use This Skill, plus 8 more sections
  • Runs Shell and Python scripts from its folder; calls kubectl and python3

What it does

K8s Debug is an agent skill from akin-ozer/cc-devops-skills. Diagnose and fix Kubernetes pods, CrashLoopBackOff, Pending, DNS, networking, storage, and rollout failures with kubectl.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `references/common_issues.md`, `references/troubleshooting_workflow.md` and `scripts/cluster_health.sh`).

It sits in DevOps & Cloud, covering Container orchestration. It works with Kubernetes. The repository describes itself as: DevOps skills for Claude Code and Codex. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Container orchestration

Example prompts

  • “/k8s-debug”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Preflight and Scope
  2. Identify the Problem Layer
  3. Gather Diagnostics with the Right Script
  4. Follow Issue-Specific Reference Workflow
  5. Apply Targeted Fixes
  6. Verify and Close

What it can do on your machine

Read from SKILL.md and the folder at commit 276af75. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Shell and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • kubectl
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

K8s Debug loads about 2.9k tokens when it runs, and up to ~7.7k if it reads all its reference files. Until then it costs about 33 tokens; SKILL.md has 895 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from akin-ozer/cc-devops-skills at commit 276af75, republished under its Apache-2.0 licence (© akin-ozer). 895 words, ~2,859 tokens.

Download SKILL.mdSave it as .claude/skills/k8s-debug/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
k8s-debug
description
Diagnose and fix Kubernetes pods, CrashLoopBackOff, Pending, DNS, networking, storage, and rollout failures with kubectl.

Kubernetes Debugging Skill

Overview

Systematic toolkit for debugging Kubernetes clusters, workloads, networking, and storage with a deterministic, safety-first workflow.

Trigger Phrases

Use this skill when requests resemble:

  • "My pod is in CrashLoopBackOff; help me find the root cause."
  • "Service DNS works in one pod but not another."
  • "Deployment rollout is stuck."
  • "Pods are Pending and not scheduling."
  • "Cluster health looks degraded after a change."
  • "PVC is pending and pods cannot mount storage."

Prerequisites

Run from the skill directory (devops-skills-plugin/skills/k8s-debug) so relative script paths work as written.

Required
  • kubectl installed and configured.
  • An active cluster context.
  • Read access to namespaces, pods, events, services, and nodes.

Quick preflight:

bash
kubectl config current-context
kubectl auth can-i get pods -A
kubectl auth can-i get events -A
kubectl get ns
  • jq for more precise filtering in ./scripts/cluster_health.sh.
  • Metrics API (metrics-server) for kubectl top.
  • In-container debug tools (nslookup, getent, curl, wget, ip) for deep network tests.

Fallback behavior:

  • If optional tools are missing, scripts continue and print warnings with reduced output.
  • If kubectl top is unavailable, continue with kubectl describe and events.

When to Use This Skill

Use this skill for:

  • Pod failures (CrashLoopBackOff, ImagePullBackOff, Pending, OOMKilled)
  • Service connectivity or DNS resolution issues
  • Network policy or ingress problems
  • Volume and storage mount failures
  • Deployment rollout issues
  • Cluster health or performance degradation
  • Resource exhaustion (CPU/memory)
  • Configuration problems (ConfigMaps, Secrets, RBAC)

Safety Rules for Disruptive Commands

Default mode is read-only diagnosis first. Only execute disruptive commands after confirming blast radius and rollback.

Commands requiring explicit confirmation:

  • kubectl delete pod ... --force --grace-period=0
  • kubectl drain ...
  • kubectl rollout restart ...
  • kubectl rollout undo ...
  • kubectl debug ... --copy-to=...

Before disruptive actions:

bash
# Snapshot current state for rollback and incident notes
kubectl get deploy,rs,pod,svc -n <namespace> -o wide
kubectl get pod <pod-name> -n <namespace> -o yaml > before-<pod-name>.yaml
kubectl get events -n <namespace> --sort-by='.lastTimestamp' > before-events.txt

Reference Navigation Map

Load only the section needed for the observed symptom.

Symptom / NeedOpenStart section
You need an end-to-end diagnosis path./references/troubleshooting_workflow.mdGeneral Debugging Workflow
Pod state is Pending, CrashLoopBackOff, or ImagePullBackOff./references/troubleshooting_workflow.mdPod Lifecycle Troubleshooting
Service reachability or DNS failure./references/troubleshooting_workflow.mdNetwork Troubleshooting Workflow
Node pressure or performance regression./references/troubleshooting_workflow.mdResource and Performance Workflow
PVC / PV / storage class issues./references/troubleshooting_workflow.mdStorage Troubleshooting Workflow
Quick symptom-to-fix lookup./references/common_issues.mdmatching issue heading
Post-mortem fix options for known issues./references/common_issues.mdSolutions sections

Scripts Overview

ScriptPurposeRequired argsOptional argsOutputFallback behavior
./scripts/cluster_health.shCluster-wide health snapshot (nodes, workloads, events, common failure states)None--strict, K8S_REQUEST_TIMEOUT env varSectioned report to stdoutContinues on check failures, tracks them in summary and exit code
./scripts/network_debug.shPod-centric network and DNS diagnostics<pod-name> (<namespace> defaults to default)--strict, --insecure, K8S_REQUEST_TIMEOUT env varSectioned report to stdoutUses secure API probe by default; insecure TLS requires explicit --insecure
./scripts/pod_diagnostics.pyDeep pod diagnostics (status, describe, YAML, events, per-container logs, node context)<pod-name>-n/--namespace, -o/--outputSectioned report to stdout or fileFails fast on missing access; skips optional metrics/log blocks with clear messages
Script Exit Codes

./scripts/cluster_health.sh and ./scripts/network_debug.sh share the same contract:

  • 0: checks completed with no check failures (warnings allowed unless --strict is set).
  • 1: one or more checks failed, or warnings occurred in --strict mode.
  • 2: blocked preconditions (for example: missing kubectl, no active context, inaccessible namespace/pod).

Deterministic Debugging Workflow

Follow this systematic approach for any Kubernetes issue:

1. Preflight and Scope
bash
kubectl config current-context
kubectl get ns
kubectl auth can-i get pods -n <namespace>

If preflight fails, stop and fix access/context first.

2. Identify the Problem Layer

Categorize the issue:

  • Application Layer: Application crashes, errors, bugs
  • Pod Layer: Pod not starting, restarting, or pending
  • Service Layer: Network connectivity, DNS issues
  • Node Layer: Node not ready, resource exhaustion
  • Cluster Layer: Control plane issues, API problems
  • Storage Layer: Volume mount failures, PVC issues
  • Configuration Layer: ConfigMap, Secret, RBAC issues
Show full SKILL.md (332 more words)Show less
3. Gather Diagnostics with the Right Script

Use the appropriate diagnostic script based on scope:

Pod-Level Diagnostics

Use ./scripts/pod_diagnostics.py for comprehensive pod analysis:

bash
python3 ./scripts/pod_diagnostics.py <pod-name> -n <namespace>

This script gathers:

  • Pod status and description
  • Pod events
  • Container logs (current and previous)
  • Resource usage
  • Node information
  • YAML configuration

Output can be saved for analysis:

bash
python3 ./scripts/pod_diagnostics.py <pod-name> -n <namespace> -o diagnostics.txt
Cluster-Level Health Check

Use ./scripts/cluster_health.sh for overall cluster diagnostics:

bash
./scripts/cluster_health.sh > cluster-health-$(date +%Y%m%d-%H%M%S).txt

This script checks:

  • Cluster info and version
  • Node status and resources
  • Pods across all namespaces
  • Failed/pending pods
  • Recent events
  • Deployments, services, statefulsets, daemonsets
  • PVCs and PVs
  • Component health
  • Common error states (CrashLoopBackOff, ImagePullBackOff)
Network Diagnostics

Use ./scripts/network_debug.sh for connectivity issues:

bash
./scripts/network_debug.sh <namespace> <pod-name>
# or force warning sensitivity / insecure TLS only when explicitly needed:
./scripts/network_debug.sh --strict <namespace> <pod-name>
./scripts/network_debug.sh --insecure <namespace> <pod-name>

This script analyzes:

  • Pod network configuration
  • DNS setup and resolution
  • Service endpoints
  • Network policies
  • Connectivity tests
  • CoreDNS logs
4. Follow Issue-Specific Reference Workflow

Based on the identified issue, consult ./references/troubleshooting_workflow.md:

  • Pod Pending: Resource/scheduling workflow
  • CrashLoopBackOff: Application crash workflow
  • ImagePullBackOff: Image pull workflow
  • Service issues: Network connectivity workflow
  • DNS failures: DNS troubleshooting workflow
  • Resource exhaustion: Performance investigation workflow
  • Storage issues: PVC binding workflow
  • Deployment stuck: Rollout workflow
5. Apply Targeted Fixes

Refer to ./references/common_issues.md for symptom-specific fixes.

6. Verify and Close

Run final verification:

bash
kubectl get pods -n <namespace> -o wide
kubectl get events -n <namespace> --sort-by='.lastTimestamp' | tail -20
kubectl rollout status deployment/<name> -n <namespace>

Issue is done when user-visible behavior is healthy and no new critical warning events appear.

Example Flows

Example 1: CrashLoopBackOff in payments Namespace
bash
python3 ./scripts/pod_diagnostics.py payments-api-7c97f95dfb-q9l7k -n payments -o payments-diagnostics.txt
kubectl logs payments-api-7c97f95dfb-q9l7k -n payments --previous --tail=100
kubectl get deploy payments-api -n payments -o yaml | grep -A 8 livenessProbe

Then open ./references/common_issues.md and apply the CrashLoopBackOff solutions.

Example 2: Service DNS/Connectivity Failure
bash
./scripts/network_debug.sh checkout checkout-api-75f49c9d8f-z6qtm
kubectl get svc checkout-api -n checkout
kubectl get endpoints checkout-api -n checkout
kubectl get networkpolicies -n checkout

Then follow Service Connectivity Workflow in ./references/troubleshooting_workflow.md.

Essential Manual Commands

Pod Debugging
bash
# View pod status
kubectl get pods -n <namespace> -o wide

# Detailed pod information
kubectl describe pod <pod-name> -n <namespace>

# View logs
kubectl logs <pod-name> -n <namespace>
kubectl logs <pod-name> -n <namespace> --previous  # Previous container
kubectl logs <pod-name> -n <namespace> -c <container>  # Specific container

# Execute commands in pod
kubectl exec <pod-name> -n <namespace> -it -- /bin/sh

# Get pod YAML
kubectl get pod <pod-name> -n <namespace> -o yaml
Service and Network Debugging
bash
# Check services
kubectl get svc -n <namespace>
kubectl describe svc <service-name> -n <namespace>

# Check endpoints
kubectl get endpoints -n <namespace>

# Test DNS
kubectl exec <pod-name> -n <namespace> -- nslookup kubernetes.default

# View events
kubectl get events -n <namespace> --sort-by='.lastTimestamp'
Resource Monitoring
bash
# Node resources
kubectl top nodes
kubectl describe nodes

# Pod resources
kubectl top pods -n <namespace>
kubectl top pod <pod-name> -n <namespace> --containers
Emergency Operations
bash
# Restart deployment
kubectl rollout restart deployment/<name> -n <namespace>

# Rollback deployment
kubectl rollout undo deployment/<name> -n <namespace>

# Force delete stuck pod
kubectl delete pod <pod-name> -n <namespace> --force --grace-period=0

# Drain node (maintenance)
kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data

# Cordon node (prevent scheduling)
kubectl cordon <node-name>

Completion Criteria

Troubleshooting session is complete when all are true:

  • Cluster context and namespace are confirmed.
  • Relevant diagnostic script output is captured.
  • Root cause is identified and tied to evidence (events/logs/config/state).
  • Any disruptive action was preceded by snapshot and rollback plan.
  • Fix verification commands show healthy state.
  • Reference path used (./references/troubleshooting_workflow.md or ./references/common_issues.md) is documented in notes.

Useful additional tools for Kubernetes debugging:

  • kubectl-debug: Advanced debugging plugin
  • stern: Multi-pod log tailing
  • kubectx/kubens: Context and namespace switching
  • k9s: Terminal UI for Kubernetes
  • lens: Desktop IDE for Kubernetes
  • Prometheus/Grafana: Monitoring and alerting
  • Jaeger/Zipkin: Distributed tracing

© akin-ozer, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in devops-skills-plugin/skills/k8s-debug of akin-ozer/cc-devops-skills.

  • SKILL.md
  • references/common_issues.md
  • references/troubleshooting_workflow.md
  • scripts/cluster_health.sh
  • scripts/network_debug.sh
  • scripts/pod_diagnostics.py
  • tests/test_pod_diagnostics.py
  • tests/test_regressions.sh

Open the folder on GitHubat commit 276af75

Compare with similar skills

K8s Debug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

K8s Debug compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
K8s Debug this skillakin-ozer/cc-devops-skills319—~2.9kAutomated safety check: PassApache-2.0
KubeSphere Multi-Tenant Managementkubesphere/kubesphere17k1 repos~3.1kAutomated safety check: PassCustom licence
Azure Diagnosticsmicrosoft/azure-skills1.5k1 repos~1.6kAutomated safety check: PassMIT
Kubeshark Installerkubeshark/kubeshark12k—~3.6kAutomated safety check: NotesApache-2.0
Sim Helmsimstudioai/sim30k—~2.2kAutomated safety check: PassApache-2.0
Helm Chart ScaffoldingCybereason-Public/owLSM28013 repos~381Automated safety check: PassGPL-2.0

Similar skills

  • Creates and queries KubeSphere users, workspaces and projects and assigns built-in roles, defaulting to least privilege and never deleting anything.

    17k GitHub starsUsed in 1 repo~3.1k tokens
    DevOps & CloudAuto-check passed
  • Azure Diagnostics

    microsoft/azure-skills

    Official

    Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage.

    1.5k GitHub starsUsed in 1 repo~1.6k tokens
    DevOps & CloudAuto-check passed
  • Kubeshark Installer

    kubeshark/kubeshark

    Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.

    12k GitHub stars~3.6k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Sim Helm

    simstudioai/sim

    Install, upgrade, and operate the Sim Helm chart on Kubernetes.

    30k GitHub stars~2.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Helm Chart Scaffolding

    Cybereason-Public/owLSM

    Comprehensive guidance for creating, organizing, and managing Helm charts for packaging and deploying Kubernetes applications.

    280 GitHub starsUsed in 13 repos~381 tokens
    DevOps & CloudAuto-check passed
  • Syntax reference for KFL2, the CEL-based display filter language used to search Kubernetes network traffic captured by Kubeshark, loaded before any filter is written.

    12k GitHub stars~3.6k tokensUpdated today
    DevOps & CloudAuto-check passed

More from akin-ozer/cc-devops-skills

All 30 skills in this repo
  • GitHub Actions Generator

    akin-ozer/cc-devops-skills

    Create, generate, or scaffold GitHub Actions workflows, action.yml, or .github/workflows CI/CD pipelines.

    319 GitHub stars~3.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Helm Generator

    akin-ozer/cc-devops-skills

    Create, scaffold, or generate Helm charts, Chart.yaml, values.yaml, templates, helpers.

    319 GitHub stars~2.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Jenkinsfile Generator

    akin-ozer/cc-devops-skills

    Generate/create/scaffold Jenkinsfile — declarative, scripted, shared library, CI/CD pipelines.

    319 GitHub stars~3.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Dockerfile Validator

    akin-ozer/cc-devops-skills

    Validate, lint, audit, or scan a Dockerfile for security and best practices.

    319 GitHub stars~2.3k tokensUpdated 2 mo ago
    Auto-check passed
  • GitHub Actions Validator

    akin-ozer/cc-devops-skills

    Validate, lint, audit, fix GitHub Actions workflows (.github/workflows).

    319 GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Jenkinsfile Validator

    akin-ozer/cc-devops-skills

    Validate, lint, audit, or check Jenkinsfiles and shared libraries.

    319 GitHub stars~2.7k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Categories

Questions about K8s Debug

What does K8s Debug do?

Diagnose and fix Kubernetes pods, CrashLoopBackOff, Pending, DNS, networking, storage, and rollout failures with kubectl. K8s Debug is an agent skill from akin-ozer/cc-devops-skills. Diagnose and fix Kubernetes pods, CrashLoopBackOff, Pending, DNS, networking, storage, and rollout failures with kubectl.

When should I use K8s Debug?

K8s Debug fits situations like: tasks that involve Container orchestration.

How do I install K8s Debug in Claude Code?

Run `npx skills add akin-ozer/cc-devops-skills --skill k8s-debug -a claude-code`. Or copy the skill folder (devops-skills-plugin/skills/k8s-debug in akin-ozer/cc-devops-skills) into .claude/skills/k8s-debug in your project. Claude Code loads it when a task matches its description.

How do I install K8s Debug in Codex?

Run `npx skills add akin-ozer/cc-devops-skills --skill k8s-debug -a codex`. Or copy the skill folder (devops-skills-plugin/skills/k8s-debug in akin-ozer/cc-devops-skills) into .agents/skills/k8s-debug in your project. Codex loads it when a task matches its description.

Can I use K8s Debug in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add akin-ozer/cc-devops-skills --skill k8s-debug -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/k8s-debug, .gemini/skills/k8s-debug, .github/skills/k8s-debug and .opencode/skills/k8s-debug in your project.

What does K8s Debug need to run?

Going by SKILL.md and its folder, K8s Debug needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (kubectl and python3). Our summary lists: Python 3; A Bash shell.

Does K8s Debug access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is K8s Debug safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does K8s Debug use?

K8s Debug is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does K8s Debug use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.8k tokens, read only when the agent opens those files.

What are the alternatives to K8s Debug?

Skills that share tags, products or a category with K8s Debug: KubeSphere Multi-Tenant Management (kubesphere/kubesphere, 17k stars), Azure Diagnostics (microsoft/azure-skills, 1.5k stars), Kubeshark Installer (kubeshark/kubeshark, 12k stars) and Sim Helm (simstudioai/sim, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains K8s Debug?

akin-ozer (a GitHub user) maintains it in akin-ozer/cc-devops-skills, which has 319 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on July 26, 2026.

Source: akin-ozer/cc-devops-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.