Official agent skill

Elasticsearch Cluster Health

by elastic in elastic/agent-skills

Diagnose a non-green Elasticsearch cluster and surface the single most likely cause with remediation.

OfficialApache-2.0Auto-check passedBackend & APIs

Install Elasticsearch Cluster Health

skills CLI
$ npx skills add elastic/agent-skills --skill elasticsearch-cluster-health -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install elastic/agent-skills elasticsearch-cluster-health --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/elastic/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/elasticsearch/elasticsearch-cluster-health .claude/skills/elasticsearch-cluster-health && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
elasticsearch-cluster-health
GitHub stars
592
Token cost
~3.3k tokens
SKILL.md length
1,247 words
Files
1
Skills in repo
26
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnose a non-green Elasticsearch cluster and surface the single most likely cause with remediation.

  • Works in 5 steps: Read the overall status. Call GET… → Localize the problem to one index. Call… → Separate trigger from root cause. Call… → …
  • An operator reports yellow
  • SKILL.md covers Environment Configuration, Process, Guidelines and Examples, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Elasticsearch Cluster Health is an agent skill from elastic/agent-skills, published by the product's own GitHub organization. Diagnose a non-green Elasticsearch cluster and surface the single most likely cause with remediation. Use when an operator reports yellow or red status, unassigned shards, allocation failures, or wants read-only triage before deeper investigation. Teaches replica-vs-primary impact, allocation decider classification, and data-loss awareness.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Elasticsearch 8.x or 9.x, self-managed or Elastic Cloud Hosted; not applicable to Elastic Cloud Serverless, where cluster, shard, and allocation APIs are…

It sits in Backend & APIs, covering Search implementation. It works with Elasticsearch. The repository describes itself as: Official Elastic Skills. The licence is Apache-2.0.

When your agent uses it

  • An operator reports yellow
  • Unassigned shards
  • Allocation failures
  • Wants read-only triage before deeper investigation

Example prompts

  • “/elasticsearch-cluster-health”

Requirements

  • Compatibility (from SKILL.md): Elasticsearch 8.x or 9.x, self-managed or Elastic Cloud Hosted; not applicable to Elastic Cloud Serverless, where cluster, shard, and allocation APIs are managed internally. Requires the `elastic` CLI ≥ 0.2 with `stack es` support.

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Read the overall status. Call GET /_cluster/health. The status field is the verdict
  2. Localize the problem to one index. Call GET /_cluster/health?level=indices and pick the index that drives the
  3. Separate trigger from root cause. Call POST /_cluster/allocation/explain with no body so Elasticsearch selects
  4. Classify the decider. Map the blocking signal to a cause class. Prefer the decider with decision: "NO" over the
  5. Recommend remediation — read-only triage ends here. Report the single most likely cause (decider class +

What it can do on your machine

Read from SKILL.md and the folder at commit baa5111. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Elasticsearch 8.x or 9.x, self-managed or Elastic Cloud Hosted; not applicable to Elastic Cloud Serverless, where cluster, shard, and allocation APIs are managed internally. Requires the `elastic` CLI ≥ 0.2 with `stack es` support.

    From compatibility in the SKILL.md frontmatter.

Context cost

Elasticsearch Cluster Health loads about 3.3k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 1,247 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from elastic/agent-skills at commit baa5111, republished under its Apache-2.0 licence (© elastic). 1,247 words, ~3,302 tokens.

Download SKILL.mdSave it as .claude/skills/elasticsearch-cluster-health/SKILL.md (or your agent's skills folder).
name
elasticsearch-cluster-health
description
Diagnose a non-green Elasticsearch cluster and surface the single most likely cause with remediation. Use when an operator reports yellow or red status, unassigned shards, allocation failures, or wants read-only triage before deeper investigation. Teaches replica-vs-primary impact, allocation decider classification, and data-loss awareness.
compatibility
Elasticsearch 8.x or 9.x, self-managed or Elastic Cloud Hosted; not applicable to Elastic Cloud Serverless, where cluster, shard, and allocation APIs are managed internally. Requires the `elastic` CLI ≥ 0.2 with `stack es` support.
metadata.author
elastic
metadata.version
0.1.0
metadata.universal
true

Diagnose Cluster Health

Triage a non-green Elasticsearch cluster read-only: localize the problem, classify the allocation decider, and report the single most likely cause with remediation. Never mutate cluster state — surface findings and let the operator act.

<!-- begin-partial: preamble -->

Environment Configuration

This skill executes Elasticsearch operations through the elastic CLI. If the elastic CLI is not installed, tell the user what it is needed for. Do not guess credentials, call the HTTP API directly, or attempt other workarounds.

This skill references operations in HTTP-shorthand form (e.g., GET /, GET /_cat/indices, GET /{index}/_mapping, GET /{index}/_settings/index.mode, POST /_query). The Operations table at the end of this document maps each shorthand to the equivalent elastic CLI command — always use the CLI rather than calling the HTTP API directly.

<!-- end-partial: preamble -->

Process

  1. Read the overall status. Call GET /_cluster/health. The status field is the verdict:

    • green — every primary and replica is assigned. Report healthy and stop.
    • yellow — every primary is assigned but at least one replica is not. Data remains readable; redundancy is degraded. This is not data loss.
    • red — at least one primary is unassigned. Data for that shard is unavailable; treat as urgent.

    Also read unassigned_shards, initializing_shards, and relocating_shards. The decision: continue only when status is yellow or red. If initializing_shards > 0 and unassigned_shards == 0, the cluster is recovering on its own — call GET /_cat/recovery to confirm progress, wait, and re-check GET /_cluster/health before escalating.

    Data needed: cluster-wide status and shard counters.

  2. Localize the problem to one index. Call GET /_cluster/health?level=indices and pick the index that drives the cluster-wide status:

    • Any red index outranks every yellow index.
    • Among reds or yellows, prefer the index with the most unassigned_shards.
    • A red system index (.security, .kibana*, .fleet-*) outranks application indices because the rest of the stack depends on it.

    Optionally call GET /_cat/shards/{index}?h=index,shard,prirep,state,unassigned.reason to list every unassigned shard on that index and see whether failures are primaries (prirep=p) or replicas (prirep=r).

    The decision: focus the next steps on exactly one index — the one whose recovery unblocks the cluster.

    Data needed: per-index status and unassigned_shards; shard role (primary vs replica) when available.

  3. Separate trigger from root cause. Call POST /_cluster/allocation/explain with no body so Elasticsearch selects an unassigned shard, or target the worst shard explicitly:

    json
    { "index": "<index>", "shard": <id>, "primary": <true|false> }

    Read these fields in order:

    • primary — false means a replica is unassigned (typical yellow); true means a primary is unassigned (typical red).
    • can_allocate — top-level allocation verdict (no, yes, throttled, no_valid_shard_copy, …).
    • unassigned_info.reason — what triggered reassignment (e.g. NODE_LEFT, INDEX_CREATED). This is not the root cause when can_allocate is no; it only explains why the shard became unassigned.
    • allocate_explanation — human-readable summary; quote it verbatim in the report.
    • node_allocation_decisions[].deciders[] — per-node decider results. Find deciders with decision: "NO"; the decider name (e.g. disk_threshold, filter, awareness) is the root cause class.

    The decision:

    • Yellow + primary: false — impact is limited to replica redundancy; no data loss. Continue to step 4 to name the blocking decider (do not stop at NODE_LEFT).
    • Red + primary: true — data for that shard is missing. Continue to step 4; if can_allocate is no_valid_shard_copy, treat as potential data loss immediately.

    Data needed: allocation-explain response for one representative unassigned shard on the chosen index.

  4. Classify the decider. Map the blocking signal to a cause class. Prefer the decider with decision: "NO" over the unassigned_info.reason trigger.

    SignalCause classTypical remediation (operator applies)
    decider: disk_threshold, decision: NODisk high/low watermark exceededFree disk on the named node, add data-node capacity, or adjust cluster.routing.allocation.disk.watermark.* after confirming usage via GET /_cat/allocation
    decider: filter or decider: awareness, decision: NOAllocation filtering or zone awarenessAdd a node that satisfies index.routing.allocation.* / awareness attributes, or adjust index/cluster allocation settings
    decider: throttling or recovery in progressTransient recoveryWait; monitor GET /_cat/recovery and re-check GET /_cluster/health
    can_allocate: no_valid_shard_copy (often with empty node_allocation_decisions)No surviving shard copySee step 5 — data loss scenario
    can_allocate: yes but shard still unassignedDelayed allocation or cluster state catch-upCheck unassigned_info.at delay; wait and re-check

    For disk pressure (common yellow scenario after NODE_LEFT): replicas relocate to remaining nodes; if a survivor is above the high watermark (cluster.routing.allocation.disk.watermark.high, default 90%), the disk_threshold decider blocks replica allocation even though primaries stay assigned. The fix is disk capacity or watermark relief — not deleting the index or forcing an empty primary.

    Data needed: decider name, explanation text, and affected node names from node_allocation_decisions.

  5. Recommend remediation — read-only triage ends here. Report the single most likely cause (decider class + verbatim allocate_explanation) and one primary remediation path. Match urgency to color and shard role.

    Yellow / replica unassigned (no data loss):

    • State clearly: all primaries are assigned; only replicas are missing; no data loss.
    • Name the real decider (e.g. disk high watermark on es-node-2), not merely “a node left”.
    • Recommend: free disk space, expand storage, add data nodes, or adjust disk watermarks after reviewing GET /_cat/allocation.
    • Do not recommend: deleting the index, allocate_empty_primary, force-allocating over a healthy primary, or restarting the entire cluster without evidence.

    Red / primary unassigned with no_valid_shard_copy (data loss risk):

    • State clearly: a primary shard is unassigned; queries/routing for that shard fail; treat as urgent and localized to the named index.
    • Explain: the only copy was on the departed node; Elasticsearch cannot allocate a primary because no valid copy exists on any remaining node (can_allocate: no_valid_shard_copy).
    • Recovery paths in order:
      1. Bring the departed node back if its data directory is intact — the shard copy returns.
      2. Restore from snapshot into the index (or a new index followed by reindex) when snapshots exist.
      3. Last resort only: POST /_cluster/reroute with allocate_empty_primary — this creates an empty primary and permanently loses all documents on that shard. State data loss explicitly; never present this as the first or casual fix.
    • Do not recommend: deleting the index without discussing data loss, or allocate_empty_primary without the data-loss warning.

    Self-healing in progress:

    • When deciders show throttling or active peer recovery, recommend waiting and re-checking read-only APIs above.

    Do not execute reroutes, snapshot restores, or settings changes — surface cause and remediation only.

Show full SKILL.md (265 more words)Show less

Guidelines

  • Read-only: Use only GET/POST explain APIs for triage. Remediation is advice; the operator performs writes.
  • Trigger ≠ cause: unassigned_info.reason: NODE_LEFT explains the event; node_allocation_decisions deciders explain why allocation still fails.
  • Replica vs primary: Yellow + primary: false = redundancy gap, not data loss. Red + primary: true = missing data for that shard.
  • One index, one cause: Pick the highest-impact index and the strongest NO decider; avoid listing every shard.
  • Cat helpers: Use GET /_cat/allocation for disk percentages per node and GET /_cat/recovery for ongoing recoveries when the decider class is unclear or recovery is in progress.

Examples

Yellow — disk watermark after node departure. Health shows yellow with unassigned replicas on logs-2025-07. Allocation explain returns primary: false, unassigned_info.reason: NODE_LEFT, but disk_threshold decider NO on es-node-2 (“above the high watermark … 90%”). Report: no data loss; root cause is disk pressure on the receiving node; remediate disk/watermark — not “node left” alone.

Red — primary with no valid copy. Health shows red on orders-2025 with one unassigned shard. Explain returns primary: true, can_allocate: no_valid_shard_copy, last_allocation_status: no_valid_shard_copy. Report: urgent; primary data missing; restore node or snapshot; mention allocate_empty_primary only as last resort with explicit data loss.

Operations

HTTP API (shorthand)elastic CLI command
GET /_cluster/healthelastic es cluster health
GET /_cluster/health?level=indiceselastic es cluster health --level indices
POST /_cluster/allocation/explainelastic es cluster allocation-explain
POST /_cluster/allocation/explain (specific shard)elastic es cluster allocation-explain --index '<index>' --shard <id> --primary true (replica: false)
GET /_cat/allocationelastic es cat allocation
GET /_cat/recoveryelastic es cat recovery
GET /_cat/shards/{index}?h=index,shard,prirep,state,unassigned.reasonelastic es cat shards --index '<index>' --h index,shard,prirep,state,unassigned.reason
POST /_cluster/reroute (last-resort empty primary — operator only)elastic es cluster reroute --commands '<json>'

© elastic, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/elasticsearch/elasticsearch-cluster-health of elastic/agent-skills.

Open the folder on GitHubat commit baa5111

Compare with similar skills

Elasticsearch Cluster Health next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Elasticsearch Cluster Health compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Elasticsearch Cluster Health this skillelastic/agent-skills592—~3.3kAutomated safety check: PassApache-2.0
Product Full-Text Searchlobehub/lobehub83k—~4.1kAutomated safety check: PassCustom licence
Foundatio Repositoriesexceptionless/Exceptionless2.5k—~1.9kAutomated safety check: PassApache-2.0
Elasticsearch Authnaspectrr/deer405—~1.2kAutomated safety check: NotesMIT
Elasticsearch Authzaspectrr/deer405—~1.8kAutomated safety check: PassMIT
Elasticsearch File Ingestaspectrr/deer405—~684Automated safety check: PassMIT

Similar skills

  • Guides work on LobeHub's own product search: the shared search repository, provider choice, Elasticsearch mappings, change syncing and reindexing.

    83k GitHub stars~4.1k tokensUpdated today
    Backend & APIsAuto-check passed
  • Foundatio Repositories

    exceptionless/Exceptionless

    Query, aggregate, patch, or paginate Exceptionless data through its Elasticsearch repository abstractions.

    2.5k GitHub stars~1.9k tokensUpdated 3 days ago
    Backend & APIsAuto-check passed
  • Elasticsearch Authn

    aspectrr/deer

    Authenticate to Elasticsearch using native, file-based, LDAP/AD, SAML, OIDC, Kerberos, JWT, or certificate realms.

    405 GitHub stars~1.2k tokensUpdated 5 mo ago
    Backend & APIsAuto-check: notes
  • Elasticsearch Authz

    aspectrr/deer

    Manage Elasticsearch RBAC: native users, roles, role mappings, document- and field-level security.

    405 GitHub stars~1.8k tokensUpdated 5 mo ago
    Backend & APIsAuto-check passed
  • Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

    405 GitHub stars~684 tokensUpdated 5 mo ago
    Backend & APIsAuto-check passed
  • Diagnose and resolve Elasticsearch security errors: 401/403 failures, TLS problems, expired API keys, role mapping mismatches, and Kibana login issues.

    405 GitHub stars~4.9k tokensUpdated 5 mo ago
    Backend & APIsAuto-check passed

More from elastic/agent-skills

All 26 skills in this repo
  • Security Alert Triage

    elastic/agent-skills

    Official

    Triage Elastic Security alerts — gather context, classify threats, create cases, and acknowledge.

    592 GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check: notes
  • Security Case Management

    elastic/agent-skills

    Official

    Create, search, update, and manage SOC cases via the Kibana Cases API.

    592 GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check: notes
  • Official

    Create, tune, and manage Elastic Security detection rules (SIEM and Endpoint).

    592 GitHub starsUsed in 1 repo~3.9k tokens
    Auto-check: notes
  • Kibana Dashboards

    elastic/agent-skills

    Official

    Create and manage Kibana Dashboards and Lens visualizations.

    592 GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Official

    Generate sample security events, attack scenarios, and synthetic alerts for Elastic Security.

    592 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check passed
  • Cloud Onboarding

    elastic/agent-skills

    Official

    Onboard an Elastic Cloud organization: configure the elastic CLI's Cloud context and API key, establish a default region, then invite users, assign predefined or custom Serverless project roles, and…

    592 GitHub stars~4.1k tokensUpdated 3 days ago
    Auto-check passed

Works with

Categories

Questions about Elasticsearch Cluster Health

What does Elasticsearch Cluster Health do?

Diagnose a non-green Elasticsearch cluster and surface the single most likely cause with remediation. Elasticsearch Cluster Health is an agent skill from elastic/agent-skills, published by the product's own GitHub organization. Diagnose a non-green Elasticsearch cluster and surface the single most likely cause with remediation.

When should I use Elasticsearch Cluster Health?

Elasticsearch Cluster Health fits situations like: an operator reports yellow; unassigned shards; allocation failures; wants read-only triage before deeper investigation.

How do I install Elasticsearch Cluster Health in Claude Code?

Run `npx skills add elastic/agent-skills --skill elasticsearch-cluster-health -a claude-code`. Or copy the skill folder (skills/elasticsearch/elasticsearch-cluster-health in elastic/agent-skills) into .claude/skills/elasticsearch-cluster-health in your project. Claude Code loads it when a task matches its description.

How do I install Elasticsearch Cluster Health in Codex?

Run `npx skills add elastic/agent-skills --skill elasticsearch-cluster-health -a codex`. Or copy the skill folder (skills/elasticsearch/elasticsearch-cluster-health in elastic/agent-skills) into .agents/skills/elasticsearch-cluster-health in your project. Codex loads it when a task matches its description.

Can I use Elasticsearch Cluster Health in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add elastic/agent-skills --skill elasticsearch-cluster-health -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/elasticsearch-cluster-health, .gemini/skills/elasticsearch-cluster-health, .github/skills/elasticsearch-cluster-health and .opencode/skills/elasticsearch-cluster-health in your project.

What does Elasticsearch Cluster Health need to run?

SKILL.md names no scripts, command-line tools or credentials: Elasticsearch Cluster Health is instructions for the agent only. Compatibility (from SKILL.md): Elasticsearch 8.x or 9.x, self-managed or Elastic Cloud Hosted; not applicable to Elastic Cloud Serverless, where cluster, shard, and allocation APIs are managed internally. Requires the `elastic` CLI ≥ 0.2 with `stack es` support..

Does Elasticsearch Cluster Health access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Elasticsearch Cluster Health safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Elasticsearch Cluster Health use?

Elasticsearch Cluster Health is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Elasticsearch Cluster Health use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Elasticsearch Cluster Health?

Skills that share tags, products or a category with Elasticsearch Cluster Health: Product Full-Text Search (lobehub/lobehub, 83k stars), Foundatio Repositories (exceptionless/Exceptionless, 2.5k stars), Elasticsearch Authn (aspectrr/deer, 405 stars) and Elasticsearch Authz (aspectrr/deer, 405 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Elasticsearch Cluster Health?

elastic (a GitHub organization, an official publisher) maintains it in elastic/agent-skills, which has 592 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on October 7, 2026.

Source: elastic/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.