Kubeshark Installer
kubeshark/kubeshark
Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.
Diagnoses failed or unhealthy CockroachDB Helm chart deployments by checking Helm release state, operator health, CrdbCluster and CrdbNode status, pod readiness, RBAC, webhooks, TLS, upgrades…
$ npx skills add cockroachdb/helm-charts --skill diagnosing-cockroachdb-helm-deployments -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install cockroachdb/helm-charts diagnosing-cockroachdb-helm-deployments --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/cockroachdb/helm-charts.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments .claude/skills/diagnosing-cockroachdb-helm-deployments && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "diagnosing-cockroachdb-helm-deployments" agent skill from https://github.com/cockroachdb/helm-charts/tree/master/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments into .claude/skills/diagnosing-cockroachdb-helm-deployments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "diagnosing-cockroachdb-helm-deployments", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/cockroachdb/helm-charts/tree/master/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deploymentsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add cockroachdb/helm-charts --skill diagnosing-cockroachdb-helm-deployments -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install cockroachdb/helm-charts diagnosing-cockroachdb-helm-deployments --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cockroachdb/helm-charts.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments .agents/skills/diagnosing-cockroachdb-helm-deployments && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "diagnosing-cockroachdb-helm-deployments" agent skill from https://github.com/cockroachdb/helm-charts/tree/master/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments into .agents/skills/diagnosing-cockroachdb-helm-deployments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "diagnosing-cockroachdb-helm-deployments", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cockroachdb/helm-charts --skill diagnosing-cockroachdb-helm-deployments -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install cockroachdb/helm-charts diagnosing-cockroachdb-helm-deployments --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cockroachdb/helm-charts.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments .cursor/skills/diagnosing-cockroachdb-helm-deployments && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "diagnosing-cockroachdb-helm-deployments" agent skill from https://github.com/cockroachdb/helm-charts/tree/master/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments into .cursor/skills/diagnosing-cockroachdb-helm-deployments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "diagnosing-cockroachdb-helm-deployments", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/cockroachdb/helm-charts.git --path skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add cockroachdb/helm-charts --skill diagnosing-cockroachdb-helm-deployments -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install cockroachdb/helm-charts diagnosing-cockroachdb-helm-deployments --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cockroachdb/helm-charts.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments .gemini/skills/diagnosing-cockroachdb-helm-deployments && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "diagnosing-cockroachdb-helm-deployments" agent skill from https://github.com/cockroachdb/helm-charts/tree/master/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments into .gemini/skills/diagnosing-cockroachdb-helm-deployments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "diagnosing-cockroachdb-helm-deployments", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install cockroachdb/helm-charts diagnosing-cockroachdb-helm-deploymentsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add cockroachdb/helm-charts --skill diagnosing-cockroachdb-helm-deployments -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/cockroachdb/helm-charts.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments .github/skills/diagnosing-cockroachdb-helm-deployments && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "diagnosing-cockroachdb-helm-deployments" agent skill from https://github.com/cockroachdb/helm-charts/tree/master/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments into .github/skills/diagnosing-cockroachdb-helm-deployments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "diagnosing-cockroachdb-helm-deployments", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add cockroachdb/helm-charts --skill diagnosing-cockroachdb-helm-deployments -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install cockroachdb/helm-charts diagnosing-cockroachdb-helm-deployments --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/cockroachdb/helm-charts.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments .opencode/skills/diagnosing-cockroachdb-helm-deployments && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "diagnosing-cockroachdb-helm-deployments" agent skill from https://github.com/cockroachdb/helm-charts/tree/master/skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments into .opencode/skills/diagnosing-cockroachdb-helm-deployments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "diagnosing-cockroachdb-helm-deployments", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
diagnosing-cockroachdb-helm-deploymentsDiagnoses failed or unhealthy CockroachDB Helm chart deployments by checking Helm release state, operator health, CrdbCluster and CrdbNode status, pod readiness, RBAC, webhooks, TLS, upgrades…
Diagnosing Cockroachdb Helm Deployments is an agent skill from cockroachdb/helm-charts. Diagnoses failed or unhealthy CockroachDB Helm chart deployments by checking Helm release state, operator health, CrdbCluster and CrdbNode status, pod readiness, RBAC, webhooks, TLS, upgrades, scaling, PVCs, DNS, and multi-region assumptions. Use when Helm install or upgrade fails, pods are not Ready, or the operator is not reconciling.
Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: CockroachDB Helm v2 charts and operator-managed crdb.cockroachlabs.com/v1beta1 resources. Requires Kubernetes read access to the operator namespace and…
It sits in DevOps & Cloud, covering Container orchestration and Deployment. It works with Helm. The repository describes itself as: Helm charts for cockroachdb. The licence is Apache-2.0.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 26e44ff. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlhelmjqopensslFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
cockroachlabs.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
CockroachDB Helm v2 charts and operator-managed crdb.cockroachlabs.com/v1beta1 resources. Requires Kubernetes read access to the operator namespace and CockroachDB namespace; some remediation requires cluster-admin or platform-team action.
From compatibility in the SKILL.md frontmatter.
Diagnosing Cockroachdb Helm Deployments loads about 6.8k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 1,621 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from cockroachdb/helm-charts at commit 26e44ff, republished under its Apache-2.0 licence (© cockroachdb). 1,621 words, ~6,774 tokens.
.claude/skills/diagnosing-cockroachdb-helm-deployments/SKILL.md (or your agent's skills folder).Diagnoses CockroachDB Helm install, upgrade, and readiness failures for operator-managed clusters. Keep the flow customer-facing: collect read-only evidence, classify the failure, and propose the smallest safe remediation. If the issue needs TSC escalation or deep operator forensics, use collecting-cockroachdb-operator-escalation-packet.
helm install or helm upgrade failsCrdbCluster.status.observedGeneration is behind metadata.generationCrdbCluster.status.reconciled is false or missing--set values usedCrdbCluster, or CrdbNode resources unless the user explicitly asks for teardown and understands data impact.CrdbCluster.status; the operator owns status.cockroach init on an existing cluster.publishNotReadyAddresses without operator-team guidance.OPERATOR_NAMESPACE, CRDB_NAMESPACE, CRDBCLUSTER, CRDB_DIAG_DIR, and the CRDB_*_PATH exports set in Step 1; opening a new shell drops those variables and later steps will hit unresolved names or read the wrong object.CrdbCluster object name from a Helm release, service name, or CockroachDB image/version string. List CrdbCluster objects and use the exact metadata.name.CrdbCluster fields, save the live CRD YAML and derive field paths from the served CRD schema for the object's apiVersion.kubectl patch, kubectl annotate, kubectl delete, kubectl scale, kubectl rollout restart, helm upgrade, drain/decommission commands, and interactive kubectl exec or kubectl debug shells.export OPERATOR_NAMESPACE="<operator-namespace>"
export OPERATOR_RELEASE="<operator-release>"
export CRDB_NAMESPACE="<cockroachdb-namespace>"
export CRDB_HELM_RELEASE="<cockroachdb-release>"
export CRDB_DIAG_DIR="${CRDB_DIAG_DIR:-$(mktemp -d)}"
# Helm release state
helm -n "$OPERATOR_NAMESPACE" status "$OPERATOR_RELEASE" || true
helm -n "$CRDB_NAMESPACE" status "$CRDB_HELM_RELEASE" || true
helm -n "$OPERATOR_NAMESPACE" history "$OPERATOR_RELEASE" || true
helm -n "$CRDB_NAMESPACE" history "$CRDB_HELM_RELEASE" || true
# Operator state
kubectl -n "$OPERATOR_NAMESPACE" get deploy,pod,svc -o wide | grep -E 'cockroach-operator|NAME'
kubectl -n "$OPERATOR_NAMESPACE" logs -l app=cockroach-operator --tail=200 || true
# CRD and CockroachDB resources
kubectl get crd crdbclusters.crdb.cockroachlabs.com crdbnodes.crdb.cockroachlabs.com
kubectl get crd crdbclusters.crdb.cockroachlabs.com -o yaml > "$CRDB_DIAG_DIR/crdbclusters-crd.yaml"
kubectl get crd crdbclusters.crdb.cockroachlabs.com -o json > "$CRDB_DIAG_DIR/crdbclusters-crd.json"
kubectl -n "$CRDB_NAMESPACE" get crdbcluster,crdbnode,pod,svc,endpoints,pvc,pdb -o wide
kubectl -n "$CRDB_NAMESPACE" get crdbcluster -o json | jq -r '
.items[]
| [.metadata.name, .apiVersion, (.metadata.labels["app.kubernetes.io/instance"] // ""), (.metadata.generation | tostring)]
| @tsv
'If no CrdbCluster rows are returned, stop the object-specific diagnosis and report that no live CrdbCluster exists in the namespace. You may collect CrdbNode owner references and labels as teardown evidence, but do not treat those values as a replacement for a discovered CrdbCluster.
If multiple CrdbCluster rows are returned, choose the target by metadata.name. Do not use the Helm release or a CockroachDB version as a substitute.
export CRDBCLUSTER="<metadata.name-from-crdbcluster-list>"
test -n "$CRDBCLUSTER"
kubectl -n "$CRDB_NAMESPACE" get crdbcluster "$CRDBCLUSTER" -o yaml > "$CRDB_DIAG_DIR/crdbcluster.yaml"
kubectl -n "$CRDB_NAMESPACE" get crdbcluster "$CRDBCLUSTER" -o json > "$CRDB_DIAG_DIR/crdbcluster.json"
kubectl -n "$CRDB_NAMESPACE" describe crdbcluster "$CRDBCLUSTER" || true
kubectl -n "$CRDB_NAMESPACE" get events --sort-by=.lastTimestamp | tail -50
export CRDBCLUSTER_API_VERSION="$(jq -r '.apiVersion | split("/")[-1]' "$CRDB_DIAG_DIR/crdbcluster.json")"
export CRDBCLUSTER_SCHEMA_JSON="$CRDB_DIAG_DIR/crdbcluster-schema.json"
jq -e --arg version "$CRDBCLUSTER_API_VERSION" '
.spec.versions[] | select(.name == $version) | .schema.openAPIV3Schema
' "$CRDB_DIAG_DIR/crdbclusters-crd.json" > "$CRDBCLUSTER_SCHEMA_JSON"
crdb_schema_has() {
jq -e --arg path "$1" '
def has_schema_path($schema; $parts):
if ($parts | length) == 0 then true
elif (($schema.properties? // {}) | has($parts[0])) then
has_schema_path($schema.properties[$parts[0]]; $parts[1:])
else false
end;
has_schema_path(.; $path | split("."))
' "$CRDBCLUSTER_SCHEMA_JSON" >/dev/null
}
crdb_first_schema_path() {
for schema_path in "$@"; do
if crdb_schema_has "$schema_path"; then
printf '%s\n' "$schema_path"
return 0
fi
done
printf '\n'
}
export CRDB_MODE_PATH="$(crdb_first_schema_path spec.mode)"
export CRDB_REGIONS_PATH="$(crdb_first_schema_path spec.regions)"
export CRDB_DESIRED_IMAGE_PATH="$(crdb_first_schema_path spec.template.spec.image spec.image.name spec.image)"
export CRDB_OBSERVED_GENERATION_PATH="$(crdb_first_schema_path status.observedGeneration)"
export CRDB_RECONCILED_PATH="$(crdb_first_schema_path status.reconciled)"
export CRDB_READY_NODES_PATH="$(crdb_first_schema_path status.readyNodes)"
export CRDB_STATUS_IMAGE_PATH="$(crdb_first_schema_path status.image status.crdbcontainerimage)"
export CRDB_STATUS_VERSION_PATH="$(crdb_first_schema_path status.version)"
export CRDB_ACTIONS_PATH="$(crdb_first_schema_path status.actions status.operatorActions)"
export CRDB_CONDITIONS_PATH="$(crdb_first_schema_path status.conditions)"
export CRDB_CERTIFICATES_PATH="$(crdb_first_schema_path spec.template.spec.certificates spec.certificates)"
printf '%s\n' \
"apiVersion=$CRDBCLUSTER_API_VERSION" \
"mode=$CRDB_MODE_PATH" \
"regions=$CRDB_REGIONS_PATH" \
"desiredImage=$CRDB_DESIRED_IMAGE_PATH" \
"observedGeneration=$CRDB_OBSERVED_GENERATION_PATH" \
"reconciled=$CRDB_RECONCILED_PATH" \
"readyNodes=$CRDB_READY_NODES_PATH" \
"statusImage=$CRDB_STATUS_IMAGE_PATH" \
"statusVersion=$CRDB_STATUS_VERSION_PATH" \
"actions=$CRDB_ACTIONS_PATH" \
"conditions=$CRDB_CONDITIONS_PATH" \
"certificates=$CRDB_CERTIFICATES_PATH"For a stuck pod or node:
kubectl -n "$CRDB_NAMESPACE" describe pod <crdb-pod>
kubectl -n "$CRDB_NAMESPACE" logs <crdb-pod> -c cockroachdb --tail=200
kubectl -n "$CRDB_NAMESPACE" logs <crdb-pod> -c cockroachdb --previous
kubectl -n "$CRDB_NAMESPACE" describe crdbnode <crdbnode-name>| Symptom | Likely Class | Next Check |
|---|---|---|
no matches for kind "CrdbCluster" | CRDs/operator not installed or not ready | CRD and operator readiness |
attempt to grant extra privileges | Helm RBAC restriction | RBAC and node-reader failures |
| TLS values validation error | TLS provider conflict | TLS and certificate failures |
| Operator pod CrashLoopBackOff or OOMKilled | Operator crash or resource limit | Operator health |
| Operator running but no reconcile logs | Watch namespace mismatch or blocked worker | Operator health |
observedGeneration behind generation | Reconcile is stuck or skipped | Reconciliation not progressing |
| Pods Pending | Scheduling, storage, topology, image pull, or node labels | Pod scheduling and storage failures |
| Pods Running but not Ready | CRDB readiness, TLS, join, DNS, network, or recovery | Pod readiness and CRDB issues |
| Upgrade stuck with mixed pod images | Version validation, rejected image, rollout dependency, scheduling | Upgrade and version validation |
| Multi-region pods cannot join | DNS, network, region list, CA mismatch | DNS, service, and network issues |
| Scale-down stuck | Decommission/drain blocked or multiple nodes decommissioning | Scale down and decommission |
| Migration labels/status stuck | Migration controller issue | debugging-cockroachdb-operator-migrations |
kubectl -n "$OPERATOR_NAMESPACE" rollout status deploy/cockroach-operator --timeout=5m
kubectl get crd crdbclusters.crdb.cockroachlabs.com -o jsonpath='{.spec.versions[*].name}{"\n"}'
kubectl get crd crdbnodes.crdb.cockroachlabs.com -o jsonpath='{.spec.versions[*].name}{"\n"}'
kubectl -n "$OPERATOR_NAMESPACE" get deploy cockroach-operator -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'Remediation:
kubectl -n "$OPERATOR_NAMESPACE" get pods -l app=cockroach-operator -o wide
kubectl -n "$OPERATOR_NAMESPACE" describe pod <operator-pod>
kubectl -n "$OPERATOR_NAMESPACE" logs -l app=cockroach-operator --tail=100
kubectl -n "$OPERATOR_NAMESPACE" get deploy cockroach-operator -o jsonpath='{.spec.template.spec.containers[0].env}{"\n"}'Interpretation:
WATCH_NAMESPACE: global mode.WATCH_NAMESPACE: the operator watches only listed namespaces.Check recent logs to see whether reconciliation is active:
kubectl -n "$OPERATOR_NAMESPACE" logs -l app=cockroach-operator --tail=100 | grep -i reconcil || trueDo not add ad hoc annotations to trigger reconciliation. If a user-approved reconcile-triggering change is required, use the chart-supported timestamp path through helm upgrade --reuse-values; this updates helm.sh/restartedAt and may roll CockroachDB pods, so treat it as a mutating operation:
helm -n "$CRDB_NAMESPACE" upgrade "$CRDB_HELM_RELEASE" <cockroachdb-chart> \
--reuse-values \
--set-string cockroachdb.crdbCluster.timestamp="$(date -u +%Y-%m-%dT%H:%M:%SZ)"If the operator is healthy but silent and no user-approved mutation is appropriate, use collecting-cockroachdb-operator-escalation-packet to gather pprof and metrics before restarting it.
Common error:
attempt to grant extra privilegesCause:
Checks:
kubectl auth can-i create clusterroles.rbac.authorization.k8s.io
kubectl auth can-i create clusterrolebindings.rbac.authorization.k8s.io
kubectl auth can-i get nodesRemediation options:
nodeReader.enabled=true and subjects matching the CockroachDB ServiceAccount.cockroachdb.crdbCluster.rbac.nodeReader.create=false only after the platform-owned binding exists.Do not set nodeReader.create=false before replacement RBAC exists.
kubectl -n "$OPERATOR_NAMESPACE" get svc cockroach-webhook-service
kubectl -n "$OPERATOR_NAMESPACE" get endpoints cockroach-webhook-service
kubectl get validatingwebhookconfigurations | grep cockroachIf webhook validation fails, verify the CA bundle:
kubectl get validatingwebhookconfiguration cockroach-webhook-config \
-o jsonpath='{.webhooks[0].clientConfig.caBundle}' | base64 -d | openssl x509 -noout -dates -subject -issuerFor scoped operators, webhook configurations may be namespace-suffixed, such as cockroach-webhook-config-<namespace>.
jq \
--arg modePath "$CRDB_MODE_PATH" \
--arg desiredImagePath "$CRDB_DESIRED_IMAGE_PATH" \
--arg observedGenerationPath "$CRDB_OBSERVED_GENERATION_PATH" \
--arg reconciledPath "$CRDB_RECONCILED_PATH" \
--arg readyNodesPath "$CRDB_READY_NODES_PATH" \
--arg statusImagePath "$CRDB_STATUS_IMAGE_PATH" \
--arg statusVersionPath "$CRDB_STATUS_VERSION_PATH" \
--arg actionsPath "$CRDB_ACTIONS_PATH" \
--arg conditionsPath "$CRDB_CONDITIONS_PATH" \
'
def value($path): if $path == "" then null else getpath($path | split(".")) end;
{
apiVersion,
name: .metadata.name,
schemaPaths: {
mode: $modePath,
desiredImage: $desiredImagePath,
observedGeneration: $observedGenerationPath,
reconciled: $reconciledPath,
readyNodes: $readyNodesPath,
statusImage: $statusImagePath,
statusVersion: $statusVersionPath,
actions: $actionsPath,
conditions: $conditionsPath
},
mode: value($modePath),
desiredImage: value($desiredImagePath),
generation: .metadata.generation,
observedGeneration: value($observedGenerationPath),
reconciled: value($reconciledPath),
readyNodes: value($readyNodesPath),
statusImage: value($statusImagePath),
statusVersion: value($statusVersionPath),
actions: value($actionsPath),
conditions: value($conditionsPath)
}' "$CRDB_DIAG_DIR/crdbcluster.json"
kubectl -n "$CRDB_NAMESPACE" get crdbnodes \
-o 'custom-columns=NAME:.metadata.name,GENERATION:.metadata.generation,OBSERVED:.status.observedGeneration,DECOMMISSION:.status.decommission,REVISION:.metadata.annotations.crdb\.cockroachlabs\.com/clusterNodeRevision,NODE_ID:.status.nodeID'CrdbNode status has no phase field. Use .status.decommission to see the current decommission state (empty for healthy nodes; draining, drained, transferringReplicas, zeroReplicas, or decommissioned otherwise) and the crdb.cockroachlabs.com/clusterNodeRevision annotation to compare the current revision against the operator's desired revision. Wrap the whole -o custom-columns=... value in single quotes so shells (especially zsh) do not glob the [] or consume the internal separators.
Checklist:
spec.mode is not Disabled.kubectl -n "$CRDB_NAMESPACE" get pods -l app.kubernetes.io/name=cockroachdb -o wide
kubectl -n "$CRDB_NAMESPACE" describe pod <crdb-pod>
kubectl -n "$CRDB_NAMESPACE" logs <crdb-pod> -c cockroachdb --tail=200
kubectl -n "$CRDB_NAMESPACE" logs <crdb-pod> -c cockroachdb --previous
kubectl -n "$CRDB_NAMESPACE" get pod <crdb-pod> -o jsonpath='{.spec.containers[0].readinessProbe}{"\n"}'
kubectl -n "$CRDB_NAMESPACE" get pods -l app.kubernetes.io/name=cockroachdb \
-o 'custom-columns=NAME:.metadata.name,IMAGE:.spec.containers[0].image,PHASE:.status.phase,READY:.status.containerStatuses[0].ready,NODE:.spec.nodeName'Common pod issues:
jq \
--arg desiredImagePath "$CRDB_DESIRED_IMAGE_PATH" \
--arg statusImagePath "$CRDB_STATUS_IMAGE_PATH" \
--arg statusVersionPath "$CRDB_STATUS_VERSION_PATH" \
--arg actionsPath "$CRDB_ACTIONS_PATH" \
--arg conditionsPath "$CRDB_CONDITIONS_PATH" \
'
def value($path): if $path == "" then null else getpath($path | split(".")) end;
{
apiVersion,
name: .metadata.name,
desiredImagePath: $desiredImagePath,
desiredImage: value($desiredImagePath),
statusImagePath: $statusImagePath,
statusImage: value($statusImagePath),
statusVersionPath: $statusVersionPath,
statusVersion: value($statusVersionPath),
actionsPath: $actionsPath,
actions: value($actionsPath),
conditionsPath: $conditionsPath,
conditions: value($conditionsPath),
annotations: .metadata.annotations
}' "$CRDB_DIAG_DIR/crdbcluster.json"
jq '.metadata.annotations' "$CRDB_DIAG_DIR/crdbcluster.json"
kubectl -n "$CRDB_NAMESPACE" get jobs
kubectl -n "$CRDB_NAMESPACE" describe job <version-checker-job>
kubectl -n "$CRDB_NAMESPACE" logs -l job-name=<version-checker-job>
kubectl -n "$CRDB_NAMESPACE" get pods -l app.kubernetes.io/name=cockroachdb \
-o 'custom-columns=NAME:.metadata.name,IMAGE:.spec.containers[0].image,REVISION:.metadata.annotations.crdb\.cockroachlabs\.com/nodeRevision,PHASE:.status.phase'Interpretation:
The operator creates separate service paths for pod DNS and join traffic. Do not change service settings without operator-team guidance.
kubectl -n "$CRDB_NAMESPACE" get service,endpoints -o wide
export CRDB_SERVICE="<sql-or-public-service-name-from-output>"
export CRDB_JOIN_SERVICE="<join-service-name-from-output>"
test -n "$CRDB_SERVICE"
test -n "$CRDB_JOIN_SERVICE"
kubectl -n "$CRDB_NAMESPACE" get service "$CRDB_SERVICE" -o yaml
kubectl -n "$CRDB_NAMESPACE" get service "$CRDB_JOIN_SERVICE" -o yaml
kubectl -n "$CRDB_NAMESPACE" get endpoints "$CRDB_SERVICE"
kubectl -n "$CRDB_NAMESPACE" get endpoints "$CRDB_JOIN_SERVICE"
kubectl -n "$CRDB_NAMESPACE" exec <crdb-pod> -c cockroachdb -- \
nslookup "$CRDB_SERVICE.$CRDB_NAMESPACE.svc.cluster.local" 2>&1 || true
kubectl -n "$CRDB_NAMESPACE" exec <crdb-pod> -c cockroachdb -- \
nslookup "$CRDB_JOIN_SERVICE.$CRDB_NAMESPACE.svc.cluster.local" 2>&1 || trueFor multi-region checks, use validating-cockroachdb-helm-multiregion.
Use configuring-cockroachdb-helm-tls for TLS mode selection and detailed certificate checks.
Quick checks:
helm template <release> <chart> -n <namespace> -f values.yaml >/tmp/rendered.yaml
kubectl -n "$CRDB_NAMESPACE" get secret,configmap | grep -E 'cockroach|crdb|cert|ca|tls'
jq --arg certificatesPath "$CRDB_CERTIFICATES_PATH" '
def value($path): if $path == "" then null else getpath($path | split(".")) end;
{certificatesPath: $certificatesPath, certificates: value($certificatesPath)}
' "$CRDB_DIAG_DIR/crdbcluster.json"
kubectl -n "$CRDB_NAMESPACE" get pod <crdb-pod> -o jsonpath='{.spec.containers[*].name}{"\n"}'
kubectl -n "$CRDB_NAMESPACE" logs <crdb-pod> -c cert-reloader --tail=100Remediation:
caProvided=true has a valid caSecret.Issuer or ClusterIssuer exists and Certificate resources become Ready.kubectl -n <namespace> describe pod <pod-name>
kubectl -n <namespace> get pvc -o wide
kubectl get storageclass
kubectl get nodes -L topology.kubernetes.io/region,topology.kubernetes.io/zone
kubectl -n <namespace> exec <crdb-pod> -c cockroachdb -- df -h /cockroach/cockroach-dataCommon causes:
storageClassName is wrong.kubectl -n "$CRDB_NAMESPACE" exec <ready-crdb-pod> -c cockroachdb -- \
/cockroach/cockroach node status --decommission
kubectl -n "$CRDB_NAMESPACE" get crdbnodes -o json | jq '[.items[] | select(.status.decommission != null and .status.decommission != "")] | {count: length, nodes: [.[] | {name: .metadata.name, decommission: .status.decommission}]}'
kubectl -n "$OPERATOR_NAMESPACE" logs -l app=cockroach-operator --tail=300 | grep -Ei 'decommission|drain|scale|blocking_ranges' || trueQuestions to answer:
Only use these after collecting evidence and confirming the risk with the customer or operator team.
Disable reconciliation for one cluster:
kubectl -n "$CRDB_NAMESPACE" patch crdbcluster "$CRDBCLUSTER" --type=merge -p '{"spec":{"mode":"Disabled"}}'
# Resume reconciliation:
kubectl -n "$CRDB_NAMESPACE" patch crdbcluster "$CRDBCLUSTER" --type=merge -p '{"spec":{"mode":"MutableOnly"}}'Restart the operator after evidence is collected:
kubectl -n "$OPERATOR_NAMESPACE" rollout restart deploy/cockroach-operatorUser-approved timestamp rolling restart:
helm -n "$CRDB_NAMESPACE" upgrade "$CRDB_HELM_RELEASE" <cockroachdb-chart> \
--reuse-values \
--set-string cockroachdb.crdbCluster.timestamp="$(date -u +%Y-%m-%dT%H:%M:%SZ)"Return findings in this order:
© cockroachdb, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments of cockroachdb/helm-charts.
Open the folder on GitHubat commit 26e44ff
Diagnosing Cockroachdb Helm Deployments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Diagnosing Cockroachdb Helm Deployments this skillcockroachdb/helm-charts | 105 | — | ~6.8k | Automated safety check: Pass | Apache-2.0 | |
| Kubeshark Installerkubeshark/kubeshark | 12k | — | ~3.6k | Automated safety check: Notes | Apache-2.0 | |
| KubeShark for KubernetesLukasNiessen/kubernetes-skill | 444 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Release Chartzabbix-community/helm-zabbix | 132 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Aks Deployment Skilltimothywarner/chatgptclass | 143 | — | ~916 | Automated safety check: Pass | Custom licence | |
| KubeSphere Gateway Managementkubesphere/kubesphere | 17k | — | ~2.9k | Automated safety check: Pass | Custom licence |
kubeshark/kubeshark
Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.
LukasNiessen/kubernetes-skill
Keeps Kubernetes manifests, Helm charts and policies grounded by diagnosing six failure modes, such as insecure defaults and API drift, and loading only matching references.
zabbix-community/helm-zabbix
Cut and publish a new release of the Zabbix Helm chart in this repository, following the versioning rules and maintainer release process documented in CONTRIBUTING.md and CLAUDE.md (bump…
timothywarner/chatgptclass
Deploy and operate workloads on Azure Kubernetes Service (AKS) the safe way.
kubesphere/kubesphere
Installs, uninstalls, checks and troubleshoots the KubeSphere Gateway extension built on ingress-nginx, including gateways stuck in bad states and Helm or pod failures.
kubesphere/kubesphere
Installs, configures, upgrades and removes KubeSphere extensions through Extension, ExtensionVersion and InstallPlan resources, including dependencies and versions.
cockroachdb/helm-charts
Collects a complete CockroachDB Operator escalation packet for TSC/TSE or operator-team handoff, including Helm state, Kubernetes resources, logs, operation-specific evidence, pprof goroutine dumps…
cockroachdb/helm-charts
Selects and validates TLS settings for CockroachDB Helm chart deployments, including self-signer, cert-manager, and external certificate modes.
cockroachdb/helm-charts
Debugs CockroachDB Operator migration scenarios, including Helm StatefulSet to v1beta1 CrdbNode migration and public operator v1alpha1 to v1beta1 migration.
cockroachdb/helm-charts
Guides customer-facing installation of CockroachDB on Kubernetes using the CockroachDB split Helm charts and operator-managed v1beta1 resources.
cockroachdb/helm-charts
Validates CockroachDB Helm chart values and Kubernetes prerequisites for operator-managed multi-region deployments.
Works with
Categories
Diagnoses failed or unhealthy CockroachDB Helm chart deployments by checking Helm release state, operator health, CrdbCluster and CrdbNode status, pod readiness, RBAC, webhooks, TLS, upgrades…. Diagnosing Cockroachdb Helm Deployments is an agent skill from cockroachdb/helm-charts. Diagnoses failed or unhealthy CockroachDB Helm chart deployments by checking Helm release state, operator health, CrdbCluster and CrdbNode status, pod readiness, RBAC, webhooks, TLS, upgrades, scaling, PVCs, DNS, and multi-region assumptions.
Diagnosing Cockroachdb Helm Deployments fits situations like: pods are not Ready; the operator is not reconciling.
Run `npx skills add cockroachdb/helm-charts --skill diagnosing-cockroachdb-helm-deployments -a claude-code`. Or copy the skill folder (skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments in cockroachdb/helm-charts) into .claude/skills/diagnosing-cockroachdb-helm-deployments in your project. Claude Code loads it when a task matches its description.
Run `npx skills add cockroachdb/helm-charts --skill diagnosing-cockroachdb-helm-deployments -a codex`. Or copy the skill folder (skills/cockroachdb-observability-and-diagnostics/diagnosing-cockroachdb-helm-deployments in cockroachdb/helm-charts) into .agents/skills/diagnosing-cockroachdb-helm-deployments in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cockroachdb/helm-charts --skill diagnosing-cockroachdb-helm-deployments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/diagnosing-cockroachdb-helm-deployments, .gemini/skills/diagnosing-cockroachdb-helm-deployments, .github/skills/diagnosing-cockroachdb-helm-deployments and .opencode/skills/diagnosing-cockroachdb-helm-deployments in your project.
Going by SKILL.md and its folder, Diagnosing Cockroachdb Helm Deployments needs the command-line tools its instructions call (kubectl, helm, jq and openssl). Compatibility (from SKILL.md): CockroachDB Helm v2 charts and operator-managed crdb.cockroachlabs.com/v1beta1 resources. Requires Kubernetes read access to the operator namespace and CockroachDB namespace; some remediation requires cluster-admin or platform-team action..
SKILL.md names 1 domain. As links in the text: cockroachlabs.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Diagnosing Cockroachdb Helm Deployments is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Diagnosing Cockroachdb Helm Deployments: Kubeshark Installer (kubeshark/kubeshark, 12k stars), KubeShark for Kubernetes (LukasNiessen/kubernetes-skill, 444 stars), Release Chart (zabbix-community/helm-zabbix, 132 stars) and Aks Deployment Skill (timothywarner/chatgptclass, 143 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
cockroachdb (a GitHub organization) maintains it in cockroachdb/helm-charts, which has 105 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 1, 2026.
Source: cockroachdb/helm-charts on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.