Kubernetes Specialist
Jeffallan/claude-skills
Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.
Full cluster health check for the home-ops Kubernetes cluster.
$ npx skills add Diaoul/home-ops --skill cluster-health -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Diaoul/home-ops cluster-health --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Diaoul/home-ops.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/cluster-health .claude/skills/cluster-health && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cluster-health" agent skill from https://github.com/Diaoul/home-ops/tree/main/.claude/skills/cluster-health into .claude/skills/cluster-health/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cluster-health", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Diaoul/home-ops/tree/main/.claude/skills/cluster-healthType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Diaoul/home-ops --skill cluster-health -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Diaoul/home-ops cluster-health --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Diaoul/home-ops.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/cluster-health .agents/skills/cluster-health && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cluster-health" agent skill from https://github.com/Diaoul/home-ops/tree/main/.claude/skills/cluster-health into .agents/skills/cluster-health/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cluster-health", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Diaoul/home-ops --skill cluster-health -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Diaoul/home-ops cluster-health --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Diaoul/home-ops.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/cluster-health .cursor/skills/cluster-health && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cluster-health" agent skill from https://github.com/Diaoul/home-ops/tree/main/.claude/skills/cluster-health into .cursor/skills/cluster-health/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cluster-health", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Diaoul/home-ops.git --path .claude/skills/cluster-health--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Diaoul/home-ops --skill cluster-health -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Diaoul/home-ops cluster-health --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Diaoul/home-ops.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/cluster-health .gemini/skills/cluster-health && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cluster-health" agent skill from https://github.com/Diaoul/home-ops/tree/main/.claude/skills/cluster-health into .gemini/skills/cluster-health/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cluster-health", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Diaoul/home-ops cluster-healthInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Diaoul/home-ops --skill cluster-health -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Diaoul/home-ops.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/cluster-health .github/skills/cluster-health && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cluster-health" agent skill from https://github.com/Diaoul/home-ops/tree/main/.claude/skills/cluster-health into .github/skills/cluster-health/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cluster-health", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Diaoul/home-ops --skill cluster-health -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Diaoul/home-ops cluster-health --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Diaoul/home-ops.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/cluster-health .opencode/skills/cluster-health && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cluster-health" agent skill from https://github.com/Diaoul/home-ops/tree/main/.claude/skills/cluster-health into .opencode/skills/cluster-health/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cluster-health", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cluster-healthFull cluster health check for the home-ops Kubernetes cluster.
Cluster Health is an agent skill from Diaoul/home-ops. Full cluster health check for the home-ops Kubernetes cluster. Checks nodes, Flux, pods, storage, certs, database, networking, security, alerts, Gatus, Victoria Logs, events, and upgrade status. Use when user asks to check cluster health, run a health check, or diagnose cluster issues.
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Container orchestration and GitOps. It works with Kubernetes. The repository describes itself as: My GitOps-managed home Kubernetes cluster... and more! :sailboat:. The licence is Unlicense.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 058d446. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectljqcurlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
alertmanager.diaoul.iostatus.diaoul.iovictoria-logs.diaoul.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cluster Health loads about 2.7k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 969 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Diaoul/home-ops at commit 058d446, republished under its Unlicense licence (© Diaoul). 969 words, ~2,667 tokens.
.claude/skills/cluster-health/SKILL.md (or your agent's skills folder).Run each check below in order. After all checks, print a summary table: area | status (✅/⚠️/❌) | one-line note. Only show ⚠️ or ❌ rows in the summary unless everything is green.
Run independent checks in parallel where possible.
kubectl get nodes -o wideFlag: any node not Ready, any pressure condition, version mismatch across nodes.
flux get all -AFlag: any resource where READY != True. Ignore SUSPENDED=True resources.
kubectl get pods -A --no-headers | grep -vE '\s(Running|Completed|Succeeded)\s'
kubectl get daemonsets -A --no-headers | awk '$2 != $4 {print}'Flag: CrashLoopBackOff, Error, OOMKilled, stuck Pending. Daemonsets where DESIRED != READY.
kubectl get cephcluster -n rook-ceph -o jsonpath='{.items[0].status.ceph.health}'
kubectl get pvc -A --no-headers | grep -v BoundFlag: anything other than HEALTH_OK. Any unbound PVCs.
for p in $(kubectl get pods -n rook-ceph -l app=rook-ceph.rbd.csi.ceph.com-nodeplugin -o name); do
n=$(kubectl get $p -n rook-ceph -o jsonpath='{.spec.nodeName}')
l=$(kubectl logs $p -c csi-rbdplugin -n rook-ceph --since=15m 2>/dev/null | grep 'directory not empty')
echo "$n ENOTEMPTY=$(echo -n "$l" | grep -c .) last=$(echo "$l" | tail -1 | awk '{print $2}')"
doneFlag: any node with a non-zero count. The window is 15m and the loop retries every ~2 min, so a live loop always shows ≥3; last= is there to confirm recency (log timestamps are UTC — compare against date -u, not local time). Don't widen the window: with a 1h lookback this check keeps reporting a node as broken for an hour after it's actually fixed. Nothing else in this skill catches it — the volume keeps working for its current pod, so pods, PVCs and Ceph health all stay green while the RBD image is silently pinned to that node forever.
What it means: NodeUnpublishVolume is failing with remove /var/lib/kubelet/pods/<uid>/volumes/kubernetes.io~csi/<pv>/mount: directory not empty and retrying every ~2 min. A leftover directory sits in the pod's mount/ dir on local disk (app wrote at the volume root while unmounted). Unpublish never completes, so NodeUnstageVolume never runs, so the image stays krbd-mapped. The damage only surfaces later, when the pod reschedules to another node and hits FailedMount ... rbd image ... is still being used.
Fix — pull the pod UID and PV from the error line, confirm nothing is mounted there, then remove the leftover dir; kubelet finishes cleanup and GCs the pod dir within ~2 min:
kubectl exec <nodeplugin-pod> -c csi-rbdplugin -n rook-ceph -- grep -c '<pod-uid>' /proc/mounts # MUST be 0
kubectl exec <nodeplugin-pod> -c csi-rbdplugin -n rook-ceph -- rmdir <mount-dir>/<leftover-dir>If a pod is already stuck mounting elsewhere, also clear the stale mapping on the old node — see [[rwo-ceph-forcedelete-hazard]] for the umount/unmap sequence. Never ceph osd blocklist add.
kubectl get pods -n miroir-system --no-headers | grep -v Running
kubectl get miroirnodes
kubectl get miroirvolumes -A --no-headers | grep -v ReadyFlag: any pod not Running. Any MiroirNode missing its default pool, showing no
CAPACITY, or with a blank DRBD version — that node's r-miroir partition is missing or
its agent has not claimed it, and it silently stops being a placement candidate. Any
MiroirVolume not Ready, or reporting fewer replicas than its StorageClass asks for
(1/1 for miroir-local, 2/2 for miroir-replicated).
kubectl get snapshotschedule.kopiur.home-operations.com -A -o json \
| jq -r '.items[]
| "\(.metadata.namespace)/\(.metadata.name) cron=\(.spec.schedule.cron) last=\((.status.lastSchedule.at // "never")[0:19]) next=\((.status.nextSchedule.at // "-")[0:19]) fails=\(.status.consecutiveFailures) suspended=\(.spec.suspend // false)"'
kubectl get snapshots.kopiur.home-operations.com -A -o json \
| jq -r '[.items[] | select(.status.phase != "Succeeded")
| "\(.metadata.namespace)/\(.metadata.name) \(.status.phase)"] | .[]'
for k in snapshotpolicy snapshotschedule restore; do
echo "$k: $(kubectl get $k.kopiur.home-operations.com -A --no-headers | wc -l)"
done
date -u +%Y-%m-%dT%H:%M:%SZFlag: any schedule that is SUSPENDED, has consecutiveFailures > 0, or whose last=
is older than the interval its own cron= implies (read the cron off the object — do not
assume a period). next= in the past by more than one interval means the scheduler has
stopped firing. Also flag any Snapshot stuck in a non-Succeeded phase across two runs —
a single Running/Deleting is just a run in flight.
The three counts should match each other (one policy, schedule and restore per stateful
app); the absolute number tracks how many apps use components/persistence, so compare
them to each other rather than to a fixed number.
Snapshots are retained on a GFS schedule, so a large total is expected — see [[kopiur-retention-design]] before diagnosing accumulation.
Transient churn is normal. Each run creates a VolumeSnapshot, an ephemeral PVC and a
mover pod, then tears them down, which produces VolumeFailedDelete (PV deleted before
its VolumeAttachment detaches), FailedScheduling (mover waiting on the ephemeral PVC)
and MissingDependency (waiting for the VolumeSnapshot to become readyToUse) in
check 17. Expect roughly one of each per schedule per run. Only treat them as a fault if
PVs are stuck Released/Failed or a schedule's consecutiveFailures is climbing.
kubectl get certificate -A -o json \
| jq -r '.items[] | "\(.metadata.namespace)/\(.metadata.name) ready=\([.status.conditions[] | select(.type=="Ready") | .status] | join("")) notAfter=\(.status.notAfter) renewal=\(.status.renewalTime)"'
date -u +%Y-%m-%dT%H:%M:%SZFlag: ready != True, or a renewal time already in the past (cert-manager should have
renewed and hasn't).
Judge by renewal, not notAfter. cert-manager renews at ~2/3 of lifetime, so a healthy
90-day cert spends weeks inside any fixed "expiring soon" window while being perfectly
fine. A CertManagerCertExpirySoon alert on a cert that is ready=True with a future
renewal is an alert-threshold problem, not a certificate problem — say so rather than
flagging the cert.
kubectl get cluster -n databaseFlag: status not Cluster in healthy state, READY < INSTANCES.
kubectl get dragonfly -A
kubectl get pods -A -l app.kubernetes.io/part-of=dragonfly -o json \
| jq -r '.items[] | select(.status.phase != "Running" or ([.status.containerStatuses[]?.restartCount] | add > 0))
| "\(.metadata.namespace)/\(.metadata.name) \(.status.phase) restarts=\([.status.containerStatuses[]?.restartCount] | add)"'Flag: any Dragonfly not Ready, fewer running pods than its REPLICAS, or any restart
count above zero.
Only the operator lives in database; the instances are created per consuming app and
follow that app's namespace, so always query all namespaces rather than a fixed list. If
the label selector returns nothing, confirm with kubectl get dragonfly -A and find the
current pod labels from one of those instances before concluding anything is down.
kubectl exec -n kube-system ds/cilium -- cilium status --brief
kubectl exec -n kube-system ds/cilium -- cilium bgp peers
kubectl get pods -n network --no-headers | grep -v RunningFlag: Cilium not OK, BGP session not established, any network pod not Running.
kubectl get gateway -n network
kubectl get httproute -A -o json | jq '[.items[] | select(.status.parents[]?.conditions[]? | select(.type=="Accepted" and .status!="True"))] | length'Flag: gateways not PROGRAMMED=True, any HTTPRoute not accepted (count > 0).
kubectl get pods -n security --no-headersFlag: Authelia or LLDAP not Running, any restarts > 0.
kubectl get pods -n observability --no-headers | grep -v RunningFlag: any pod not Running.
curl -s https://alertmanager.diaoul.io/api/v2/alerts \
| jq '[.[] | select(.status.silencedBy == [] and (.labels.alertname | test("InfoInhibitor|Watchdog") | not))] | .[] | {alert: .labels.alertname, severity: .labels.severity, namespace: .labels.namespace}'Flag: any critical alerts. Warning alerts note but don't fail.
curl -s https://status.diaoul.io/api/v1/endpoints/statuses \
| jq '[.[] | select(.results[-1].success == false) | .name]'Flag: any service in the list (non-empty = failing probes).
curl -s 'https://victoria-logs.diaoul.io/select/logsql/query?start=1h&query=%28level%3AERROR%20OR%20level%3AWARNING%29%20%7C%20stats%20by%20%28kubernetes.pod_namespace%29%20count%28%29%20as%20cnt'The query must be URL-encoded into the query string as above: the sandbox rejects
curl --data-urlencode (it reads as a POST), and an empty query arg returns
`query` arg cannot be empty.
Flag: namespaces with unusually high error counts. Judge relatively, not against a fixed
threshold — kube-system carries constant background noise, and rook-ceph spikes for
an hour after any Ceph or node-level change. Compare namespaces against each other and
against what the rest of this run already found: a spike in a namespace whose pods,
PVCs and Flux resources are all green is usually the tail of something that already
resolved. An app namespace that is normally silent appearing at all is the real signal.
kubectl get events -A --field-selector=type=Warning --sort-by='.lastTimestamp' | tail -20Flag: OOMKilled, FailedScheduling, recurring BackOff.
kubectl top nodes
kubectl top pods -A --sort-by=memory | head -20Flag: any node >85% memory, any node >90% CPU sustained.
kubectl get talosupgrade,kubernetesupgrade -AFlag: Phase not Completed — upgrade in progress or failed.
© Diaoul, Unlicense. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/cluster-health of Diaoul/home-ops.
Open the folder on GitHubat commit 058d446
Cluster Health next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cluster Health this skillDiaoul/home-ops | 120 | — | ~2.7k | Automated safety check: Pass | Unlicense | |
| Kubernetes SpecialistJeffallan/claude-skills | 12k | 1 repos | ~2.1k | Automated safety check: Pass | MIT | |
| Kubernetes ArchitectCybereason-Public/owLSM | 280 | 9 repos | ~2.6k | Automated safety check: Pass | GPL-2.0 | |
| Devopsnicepkg/auto-company | 195 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| GitOps with ArgoCD and Fluxwshobson/agents | 40k | 12 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Signozqjoly/GitOps | 112 | — | ~6.1k | Automated safety check: Pass | WTFPL |
Jeffallan/claude-skills
Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.
Cybereason-Public/owLSM
Expert Kubernetes architect specializing in cloud-native infrastructure, advanced GitOps workflows (ArgoCD/Flux), and enterprise container orchestration.
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
wshobson/agents
Sets up GitOps continuous delivery for Kubernetes with ArgoCD or Flux, covering installation, repository layout, sync policies, progressive delivery and secrets.
qjoly/GitOps
Manage the self-hosted SigNoz observability stack in this GitOps repo.
Yikai-Liao/symusic
Creates Dockerfiles, configures CI/CD pipelines, writes Kubernetes manifests, and generates Terraform/Pulumi infrastructure templates.
Works with
Categories
Full cluster health check for the home-ops Kubernetes cluster. Cluster Health is an agent skill from Diaoul/home-ops. Full cluster health check for the home-ops Kubernetes cluster.
Cluster Health fits situations like: user asks to check cluster health; run a health check; diagnose cluster issues.
Run `npx skills add Diaoul/home-ops --skill cluster-health -a claude-code`. Or copy the skill folder (.claude/skills/cluster-health in Diaoul/home-ops) into .claude/skills/cluster-health in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Diaoul/home-ops --skill cluster-health -a codex`. Or copy the skill folder (.claude/skills/cluster-health in Diaoul/home-ops) into .agents/skills/cluster-health in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Diaoul/home-ops --skill cluster-health -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cluster-health, .gemini/skills/cluster-health, .github/skills/cluster-health and .opencode/skills/cluster-health in your project.
Going by SKILL.md and its folder, Cluster Health needs the command-line tools its instructions call (kubectl, jq and curl).
SKILL.md names 3 domains. In commands or code: alertmanager.diaoul.io, status.diaoul.io and victoria-logs.diaoul.io; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Cluster Health is published under the Unlicense licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Cluster Health: Kubernetes Specialist (Jeffallan/claude-skills, 12k stars), Kubernetes Architect (Cybereason-Public/owLSM, 280 stars), Devops (nicepkg/auto-company, 195 stars) and GitOps with ArgoCD and Flux (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Diaoul (a GitHub user) maintains it in Diaoul/home-ops, which has 120 GitHub stars. The repository was last updated on October 11, 2026.
Source: Diaoul/home-ops on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.