Agent skill

Cluster Health

by Diaoul in Diaoul/home-ops

Full cluster health check for the home-ops Kubernetes cluster.

UnlicenseAuto-check passedDevOps & Cloud

Install Cluster Health

skills CLI
$ npx skills add Diaoul/home-ops --skill cluster-health -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Diaoul/home-ops cluster-health --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Diaoul/home-ops.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/cluster-health .claude/skills/cluster-health && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cluster-health
GitHub stars
120
Token cost
~2.7k tokens
SKILL.md length
969 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
Unlicense

At a glance

Full cluster health check for the home-ops Kubernetes cluster.

  • Works in 12 steps: Nodes → Flux → Pods → …
  • User asks to check cluster health
  • SKILL.md covers 1. Nodes, 2. Flux, 3. Pods and 4. Rook-Ceph, plus 15 more sections
  • Calls kubectl, jq and curl; reaches alertmanager.diaoul.io and status.diaoul.io

What it does

Cluster Health is an agent skill from Diaoul/home-ops. Full cluster health check for the home-ops Kubernetes cluster. Checks nodes, Flux, pods, storage, certs, database, networking, security, alerts, Gatus, Victoria Logs, events, and upgrade status. Use when user asks to check cluster health, run a health check, or diagnose cluster issues.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Container orchestration and GitOps. It works with Kubernetes. The repository describes itself as: My GitOps-managed home Kubernetes cluster... and more! :sailboat:. The licence is Unlicense.

When your agent uses it

  • User asks to check cluster health
  • Run a health check
  • Diagnose cluster issues

Example prompts

  • “/cluster-health”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Nodes
  2. Flux
  3. Pods
  4. Rook-Ceph
  5. miroir
  6. Kopiur Backups
  7. Certificates
  8. CloudNative-PG
  9. Dragonfly
  10. Networking
  11. HTTPRoutes / Gateways
  12. Security

What it can do on your machine

Read from SKILL.md and the folder at commit 058d446. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • jq
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • alertmanager.diaoul.io
    • status.diaoul.io
    • victoria-logs.diaoul.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cluster Health loads about 2.7k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 969 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Diaoul/home-ops at commit 058d446, republished under its Unlicense licence (© Diaoul). 969 words, ~2,667 tokens.

Download SKILL.mdSave it as .claude/skills/cluster-health/SKILL.md (or your agent's skills folder).
name
cluster-health
description
Full cluster health check for the home-ops Kubernetes cluster. Checks nodes, Flux, pods, storage, certs, database, networking, security, alerts, Gatus, Victoria Logs, events, and upgrade status. Use when user asks to check cluster health, run a health check, or diagnose cluster issues.

Run each check below in order. After all checks, print a summary table: area | status (✅/⚠️/❌) | one-line note. Only show ⚠️ or ❌ rows in the summary unless everything is green.

Run independent checks in parallel where possible.

1. Nodes

sh
kubectl get nodes -o wide

Flag: any node not Ready, any pressure condition, version mismatch across nodes.

2. Flux

sh
flux get all -A

Flag: any resource where READY != True. Ignore SUSPENDED=True resources.

3. Pods

sh
kubectl get pods -A --no-headers | grep -vE '\s(Running|Completed|Succeeded)\s'
kubectl get daemonsets -A --no-headers | awk '$2 != $4 {print}'

Flag: CrashLoopBackOff, Error, OOMKilled, stuck Pending. Daemonsets where DESIRED != READY.

4. Rook-Ceph

sh
kubectl get cephcluster -n rook-ceph -o jsonpath='{.items[0].status.ceph.health}'
kubectl get pvc -A --no-headers | grep -v Bound

Flag: anything other than HEALTH_OK. Any unbound PVCs.

4b. Stuck CSI unpublish (silent stale RBD mappings)
sh
for p in $(kubectl get pods -n rook-ceph -l app=rook-ceph.rbd.csi.ceph.com-nodeplugin -o name); do
  n=$(kubectl get $p -n rook-ceph -o jsonpath='{.spec.nodeName}')
  l=$(kubectl logs $p -c csi-rbdplugin -n rook-ceph --since=15m 2>/dev/null | grep 'directory not empty')
  echo "$n ENOTEMPTY=$(echo -n "$l" | grep -c .) last=$(echo "$l" | tail -1 | awk '{print $2}')"
done

Flag: any node with a non-zero count. The window is 15m and the loop retries every ~2 min, so a live loop always shows ≥3; last= is there to confirm recency (log timestamps are UTC — compare against date -u, not local time). Don't widen the window: with a 1h lookback this check keeps reporting a node as broken for an hour after it's actually fixed. Nothing else in this skill catches it — the volume keeps working for its current pod, so pods, PVCs and Ceph health all stay green while the RBD image is silently pinned to that node forever.

What it means: NodeUnpublishVolume is failing with remove /var/lib/kubelet/pods/<uid>/volumes/kubernetes.io~csi/<pv>/mount: directory not empty and retrying every ~2 min. A leftover directory sits in the pod's mount/ dir on local disk (app wrote at the volume root while unmounted). Unpublish never completes, so NodeUnstageVolume never runs, so the image stays krbd-mapped. The damage only surfaces later, when the pod reschedules to another node and hits FailedMount ... rbd image ... is still being used.

Fix — pull the pod UID and PV from the error line, confirm nothing is mounted there, then remove the leftover dir; kubelet finishes cleanup and GCs the pod dir within ~2 min:

sh
kubectl exec <nodeplugin-pod> -c csi-rbdplugin -n rook-ceph -- grep -c '<pod-uid>' /proc/mounts   # MUST be 0
kubectl exec <nodeplugin-pod> -c csi-rbdplugin -n rook-ceph -- rmdir <mount-dir>/<leftover-dir>

If a pod is already stuck mounting elsewhere, also clear the stale mapping on the old node — see [[rwo-ceph-forcedelete-hazard]] for the umount/unmap sequence. Never ceph osd blocklist add.

5. miroir

sh
kubectl get pods -n miroir-system --no-headers | grep -v Running
kubectl get miroirnodes
kubectl get miroirvolumes -A --no-headers | grep -v Ready

Flag: any pod not Running. Any MiroirNode missing its default pool, showing no CAPACITY, or with a blank DRBD version — that node's r-miroir partition is missing or its agent has not claimed it, and it silently stops being a placement candidate. Any MiroirVolume not Ready, or reporting fewer replicas than its StorageClass asks for (1/1 for miroir-local, 2/2 for miroir-replicated).

6. Kopiur Backups

sh
kubectl get snapshotschedule.kopiur.home-operations.com -A -o json \
  | jq -r '.items[]
      | "\(.metadata.namespace)/\(.metadata.name) cron=\(.spec.schedule.cron) last=\((.status.lastSchedule.at // "never")[0:19]) next=\((.status.nextSchedule.at // "-")[0:19]) fails=\(.status.consecutiveFailures) suspended=\(.spec.suspend // false)"'
kubectl get snapshots.kopiur.home-operations.com -A -o json \
  | jq -r '[.items[] | select(.status.phase != "Succeeded")
            | "\(.metadata.namespace)/\(.metadata.name) \(.status.phase)"] | .[]'
for k in snapshotpolicy snapshotschedule restore; do
  echo "$k: $(kubectl get $k.kopiur.home-operations.com -A --no-headers | wc -l)"
done
date -u +%Y-%m-%dT%H:%M:%SZ

Flag: any schedule that is SUSPENDED, has consecutiveFailures > 0, or whose last= is older than the interval its own cron= implies (read the cron off the object — do not assume a period). next= in the past by more than one interval means the scheduler has stopped firing. Also flag any Snapshot stuck in a non-Succeeded phase across two runs — a single Running/Deleting is just a run in flight.

The three counts should match each other (one policy, schedule and restore per stateful app); the absolute number tracks how many apps use components/persistence, so compare them to each other rather than to a fixed number.

Snapshots are retained on a GFS schedule, so a large total is expected — see [[kopiur-retention-design]] before diagnosing accumulation.

Transient churn is normal. Each run creates a VolumeSnapshot, an ephemeral PVC and a mover pod, then tears them down, which produces VolumeFailedDelete (PV deleted before its VolumeAttachment detaches), FailedScheduling (mover waiting on the ephemeral PVC) and MissingDependency (waiting for the VolumeSnapshot to become readyToUse) in check 17. Expect roughly one of each per schedule per run. Only treat them as a fault if PVs are stuck Released/Failed or a schedule's consecutiveFailures is climbing.

Show full SKILL.md (388 more words)Show less

7. Certificates

sh
kubectl get certificate -A -o json \
  | jq -r '.items[] | "\(.metadata.namespace)/\(.metadata.name) ready=\([.status.conditions[] | select(.type=="Ready") | .status] | join("")) notAfter=\(.status.notAfter) renewal=\(.status.renewalTime)"'
date -u +%Y-%m-%dT%H:%M:%SZ

Flag: ready != True, or a renewal time already in the past (cert-manager should have renewed and hasn't).

Judge by renewal, not notAfter. cert-manager renews at ~2/3 of lifetime, so a healthy 90-day cert spends weeks inside any fixed "expiring soon" window while being perfectly fine. A CertManagerCertExpirySoon alert on a cert that is ready=True with a future renewal is an alert-threshold problem, not a certificate problem — say so rather than flagging the cert.

8. CloudNative-PG

sh
kubectl get cluster -n database

Flag: status not Cluster in healthy state, READY < INSTANCES.

9. Dragonfly

sh
kubectl get dragonfly -A
kubectl get pods -A -l app.kubernetes.io/part-of=dragonfly -o json \
  | jq -r '.items[] | select(.status.phase != "Running" or ([.status.containerStatuses[]?.restartCount] | add > 0))
      | "\(.metadata.namespace)/\(.metadata.name) \(.status.phase) restarts=\([.status.containerStatuses[]?.restartCount] | add)"'

Flag: any Dragonfly not Ready, fewer running pods than its REPLICAS, or any restart count above zero.

Only the operator lives in database; the instances are created per consuming app and follow that app's namespace, so always query all namespaces rather than a fixed list. If the label selector returns nothing, confirm with kubectl get dragonfly -A and find the current pod labels from one of those instances before concluding anything is down.

10. Networking

sh
kubectl exec -n kube-system ds/cilium -- cilium status --brief
kubectl exec -n kube-system ds/cilium -- cilium bgp peers
kubectl get pods -n network --no-headers | grep -v Running

Flag: Cilium not OK, BGP session not established, any network pod not Running.

11. HTTPRoutes / Gateways

sh
kubectl get gateway -n network
kubectl get httproute -A -o json | jq '[.items[] | select(.status.parents[]?.conditions[]? | select(.type=="Accepted" and .status!="True"))] | length'

Flag: gateways not PROGRAMMED=True, any HTTPRoute not accepted (count > 0).

12. Security

sh
kubectl get pods -n security --no-headers

Flag: Authelia or LLDAP not Running, any restarts > 0.

13. Observability Stack

sh
kubectl get pods -n observability --no-headers | grep -v Running

Flag: any pod not Running.

14. Firing Alerts

sh
curl -s https://alertmanager.diaoul.io/api/v2/alerts \
  | jq '[.[] | select(.status.silencedBy == [] and (.labels.alertname | test("InfoInhibitor|Watchdog") | not))] | .[] | {alert: .labels.alertname, severity: .labels.severity, namespace: .labels.namespace}'

Flag: any critical alerts. Warning alerts note but don't fail.

15. Gatus — Per-Service Status

sh
curl -s https://status.diaoul.io/api/v1/endpoints/statuses \
  | jq '[.[] | select(.results[-1].success == false) | .name]'

Flag: any service in the list (non-empty = failing probes).

16. Victoria Logs — Error/Warning Scan

sh
curl -s 'https://victoria-logs.diaoul.io/select/logsql/query?start=1h&query=%28level%3AERROR%20OR%20level%3AWARNING%29%20%7C%20stats%20by%20%28kubernetes.pod_namespace%29%20count%28%29%20as%20cnt'

The query must be URL-encoded into the query string as above: the sandbox rejects curl --data-urlencode (it reads as a POST), and an empty query arg returns `query` arg cannot be empty.

Flag: namespaces with unusually high error counts. Judge relatively, not against a fixed threshold — kube-system carries constant background noise, and rook-ceph spikes for an hour after any Ceph or node-level change. Compare namespaces against each other and against what the rest of this run already found: a spike in a namespace whose pods, PVCs and Flux resources are all green is usually the tail of something that already resolved. An app namespace that is normally silent appearing at all is the real signal.

17. Kubernetes Warning Events

sh
kubectl get events -A --field-selector=type=Warning --sort-by='.lastTimestamp' | tail -20

Flag: OOMKilled, FailedScheduling, recurring BackOff.

18. Resource Pressure

sh
kubectl top nodes
kubectl top pods -A --sort-by=memory | head -20

Flag: any node >85% memory, any node >90% CPU sustained.

19. System Upgrades

sh
kubectl get talosupgrade,kubernetesupgrade -A

Flag: Phase not Completed — upgrade in progress or failed.

© Diaoul, Unlicense. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/cluster-health of Diaoul/home-ops.

Open the folder on GitHubat commit 058d446

Compare with similar skills

Cluster Health next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cluster Health compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cluster Health this skillDiaoul/home-ops120—~2.7kAutomated safety check: PassUnlicense
Kubernetes SpecialistJeffallan/claude-skills12k1 repos~2.1kAutomated safety check: PassMIT
Kubernetes ArchitectCybereason-Public/owLSM2809 repos~2.6kAutomated safety check: PassGPL-2.0
Devopsnicepkg/auto-company1952 repos~814Automated safety check: PassMIT
GitOps with ArgoCD and Fluxwshobson/agents40k12 repos~1.5kAutomated safety check: PassMIT
Signozqjoly/GitOps112—~6.1kAutomated safety check: PassWTFPL

Similar skills

  • Kubernetes Specialist

    Jeffallan/claude-skills

    Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    DevOps & CloudAuto-check passed
  • Kubernetes Architect

    Cybereason-Public/owLSM

    Expert Kubernetes architect specializing in cloud-native infrastructure, advanced GitOps workflows (ArgoCD/Flux), and enterprise container orchestration.

    280 GitHub starsUsed in 9 repos~2.6k tokens
    DevOps & CloudAuto-check passed
  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    195 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • Sets up GitOps continuous delivery for Kubernetes with ArgoCD or Flux, covering installation, repository layout, sync policies, progressive delivery and secrets.

    40k GitHub starsUsed in 12 repos~1.5k tokens
    DevOps & CloudAuto-check passed
  • Signoz

    qjoly/GitOps

    Manage the self-hosted SigNoz observability stack in this GitOps repo.

    112 GitHub stars~6.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Devops Engineer

    Yikai-Liao/symusic

    Creates Dockerfiles, configures CI/CD pipelines, writes Kubernetes manifests, and generates Terraform/Pulumi infrastructure templates.

    189 GitHub starsUsed in 1 repo~1.5k tokens
    DevOps & CloudAuto-check passed

Works with

Categories

Questions about Cluster Health

What does Cluster Health do?

Full cluster health check for the home-ops Kubernetes cluster. Cluster Health is an agent skill from Diaoul/home-ops. Full cluster health check for the home-ops Kubernetes cluster.

When should I use Cluster Health?

Cluster Health fits situations like: user asks to check cluster health; run a health check; diagnose cluster issues.

How do I install Cluster Health in Claude Code?

Run `npx skills add Diaoul/home-ops --skill cluster-health -a claude-code`. Or copy the skill folder (.claude/skills/cluster-health in Diaoul/home-ops) into .claude/skills/cluster-health in your project. Claude Code loads it when a task matches its description.

How do I install Cluster Health in Codex?

Run `npx skills add Diaoul/home-ops --skill cluster-health -a codex`. Or copy the skill folder (.claude/skills/cluster-health in Diaoul/home-ops) into .agents/skills/cluster-health in your project. Codex loads it when a task matches its description.

Can I use Cluster Health in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Diaoul/home-ops --skill cluster-health -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cluster-health, .gemini/skills/cluster-health, .github/skills/cluster-health and .opencode/skills/cluster-health in your project.

What does Cluster Health need to run?

Going by SKILL.md and its folder, Cluster Health needs the command-line tools its instructions call (kubectl, jq and curl).

Does Cluster Health access the network?

SKILL.md names 3 domains. In commands or code: alertmanager.diaoul.io, status.diaoul.io and victoria-logs.diaoul.io; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Cluster Health safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cluster Health use?

Cluster Health is published under the Unlicense licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cluster Health use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cluster Health?

Skills that share tags, products or a category with Cluster Health: Kubernetes Specialist (Jeffallan/claude-skills, 12k stars), Kubernetes Architect (Cybereason-Public/owLSM, 280 stars), Devops (nicepkg/auto-company, 195 stars) and GitOps with ArgoCD and Flux (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cluster Health?

Diaoul (a GitHub user) maintains it in Diaoul/home-ops, which has 120 GitHub stars. The repository was last updated on October 11, 2026.

Source: Diaoul/home-ops on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.