Elasticsearch Cluster Health

作者 elasticbaa511126ba2无许可证592 个星标收录于 2026年10月8日更新于 2026年10月8日仓库昨天更新

Diagnose a non-green Elasticsearch cluster and surface the single most likely cause with remediation. Use when an operator reports yellow or red status, unassigned shards, allocation failures, or wants read-only triage before deeper investigation. Teaches replica-vs-primary impact, allocation decider classification, and data-loss awareness.

仅含说明DevOps & Cloud
AI 生成的概览

对非绿色 Elasticsearch 集群进行只读排查,指出最可能的原因与修复建议。

功能
指导运维人员对黄色或红色 Elasticsearch 集群进行只读诊断:读取集群健康状态,将问题定位到单个索引,并利用分配解释输出对阻塞的分配决策器进行分类。它区分副本缺失与主分片缺失,提示可能的数据丢失,并报告最可能的原因及一条修复路径。它不会更改集群状态,修复仅为建议。
适用场景
适用于运维人员报告集群状态为黄色或红色、存在未分配分片或分配失败,并希望在深入调查前进行只读排查的场景。也适合判断问题是冗余缺口还是实际数据丢失。
运行要求
Elasticsearch 8.x 或 9.x,自管理或 Elastic Cloud Hosted(不适用于 Elastic Cloud Serverless)。需要 0.2 或更高版本且支持 stack es 的 elastic CLI,并已配置凭据。仅为说明文档,不附带脚本。

Diagnose Cluster Health

Triage a non-green Elasticsearch cluster read-only: localize the problem, classify the allocation decider, and report the single most likely cause with remediation. Never mutate cluster state — surface findings and let the operator act.

<!-- begin-partial: preamble -->

Environment Configuration

This skill executes Elasticsearch operations through the elastic CLI. If the elastic CLI is not installed, tell the user what it is needed for. Do not guess credentials, call the HTTP API directly, or attempt other workarounds.

This skill references operations in HTTP-shorthand form (e.g., GET /, GET /_cat/indices, GET /{index}/_mapping, GET /{index}/_settings/index.mode, POST /_query). The Operations table at the end of this document maps each shorthand to the equivalent elastic CLI command — always use the CLI rather than calling the HTTP API directly.

<!-- end-partial: preamble -->

Process

  1. Read the overall status. Call GET /_cluster/health. The status field is the verdict:

    • green — every primary and replica is assigned. Report healthy and stop.
    • yellow — every primary is assigned but at least one replica is not. Data remains readable; redundancy is degraded. This is not data loss.
    • red — at least one primary is unassigned. Data for that shard is unavailable; treat as urgent.

    Also read unassigned_shards, initializing_shards, and relocating_shards. The decision: continue only when status is yellow or red. If initializing_shards > 0 and unassigned_shards == 0, the cluster is recovering on its own — call GET /_cat/recovery to confirm progress, wait, and re-check GET /_cluster/health before escalating.

    Data needed: cluster-wide status and shard counters.

  2. Localize the problem to one index. Call GET /_cluster/health?level=indices and pick the index that drives the cluster-wide status:

    • Any red index outranks every yellow index.
    • Among reds or yellows, prefer the index with the most unassigned_shards.
    • A red system index (.security, .kibana*, .fleet-*) outranks application indices because the rest of the stack depends on it.

    Optionally call GET /_cat/shards/{index}?h=index,shard,prirep,state,unassigned.reason to list every unassigned shard on that index and see whether failures are primaries (prirep=p) or replicas (prirep=r).

    The decision: focus the next steps on exactly one index — the one whose recovery unblocks the cluster.

    Data needed: per-index status and unassigned_shards; shard role (primary vs replica) when available.

  3. Separate trigger from root cause. Call POST /_cluster/allocation/explain with no body so Elasticsearch selects an unassigned shard, or target the worst shard explicitly:

    json
    { "index": "<index>", "shard": <id>, "primary": <true|false> }

    Read these fields in order:

    • primary — false means a replica is unassigned (typical yellow); true means a primary is unassigned (typical red).
    • can_allocate — top-level allocation verdict (no, yes, throttled, no_valid_shard_copy, …).
    • unassigned_info.reason — what triggered reassignment (e.g. NODE_LEFT, INDEX_CREATED). This is not the root cause when can_allocate is no; it only explains why the shard became unassigned.
    • allocate_explanation — human-readable summary; quote it verbatim in the report.
    • node_allocation_decisions[].deciders[] — per-node decider results. Find deciders with decision: "NO"; the decider name (e.g. disk_threshold, filter, awareness) is the root cause class.

    The decision:

    • Yellow + primary: false — impact is limited to replica redundancy; no data loss. Continue to step 4 to name the blocking decider (do not stop at NODE_LEFT).
    • Red + primary: true — data for that shard is missing. Continue to step 4; if can_allocate is no_valid_shard_copy, treat as potential data loss immediately.

    Data needed: allocation-explain response for one representative unassigned shard on the chosen index.

  4. Classify the decider. Map the blocking signal to a cause class. Prefer the decider with decision: "NO" over the unassigned_info.reason trigger.

    SignalCause classTypical remediation (operator applies)
    decider: disk_threshold, decision: NODisk high/low watermark exceededFree disk on the named node, add data-node capacity, or adjust cluster.routing.allocation.disk.watermark.* after confirming usage via GET /_cat/allocation
    decider: filter or decider: awareness, decision: NOAllocation filtering or zone awarenessAdd a node that satisfies index.routing.allocation.* / awareness attributes, or adjust index/cluster allocation settings
    decider: throttling or recovery in progressTransient recoveryWait; monitor GET /_cat/recovery and re-check GET /_cluster/health
    can_allocate: no_valid_shard_copy (often with empty node_allocation_decisions)No surviving shard copySee step 5 — data loss scenario
    can_allocate: yes but shard still unassignedDelayed allocation or cluster state catch-upCheck unassigned_info.at delay; wait and re-check

    For disk pressure (common yellow scenario after NODE_LEFT): replicas relocate to remaining nodes; if a survivor is above the high watermark (cluster.routing.allocation.disk.watermark.high, default 90%), the disk_threshold decider blocks replica allocation even though primaries stay assigned. The fix is disk capacity or watermark relief — not deleting the index or forcing an empty primary.

    Data needed: decider name, explanation text, and affected node names from node_allocation_decisions.

  5. Recommend remediation — read-only triage ends here. Report the single most likely cause (decider class + verbatim allocate_explanation) and one primary remediation path. Match urgency to color and shard role.

    Yellow / replica unassigned (no data loss):

    • State clearly: all primaries are assigned; only replicas are missing; no data loss.
    • Name the real decider (e.g. disk high watermark on es-node-2), not merely “a node left”.
    • Recommend: free disk space, expand storage, add data nodes, or adjust disk watermarks after reviewing GET /_cat/allocation.
    • Do not recommend: deleting the index, allocate_empty_primary, force-allocating over a healthy primary, or restarting the entire cluster without evidence.

    Red / primary unassigned with no_valid_shard_copy (data loss risk):

    • State clearly: a primary shard is unassigned; queries/routing for that shard fail; treat as urgent and localized to the named index.
    • Explain: the only copy was on the departed node; Elasticsearch cannot allocate a primary because no valid copy exists on any remaining node (can_allocate: no_valid_shard_copy).
    • Recovery paths in order:
      1. Bring the departed node back if its data directory is intact — the shard copy returns.
      2. Restore from snapshot into the index (or a new index followed by reindex) when snapshots exist.
      3. Last resort only: POST /_cluster/reroute with allocate_empty_primary — this creates an empty primary and permanently loses all documents on that shard. State data loss explicitly; never present this as the first or casual fix.
    • Do not recommend: deleting the index without discussing data loss, or allocate_empty_primary without the data-loss warning.

    Self-healing in progress:

    • When deciders show throttling or active peer recovery, recommend waiting and re-checking read-only APIs above.

    Do not execute reroutes, snapshot restores, or settings changes — surface cause and remediation only.

Guidelines

  • Read-only: Use only GET/POST explain APIs for triage. Remediation is advice; the operator performs writes.
  • Trigger ≠ cause: unassigned_info.reason: NODE_LEFT explains the event; node_allocation_decisions deciders explain why allocation still fails.
  • Replica vs primary: Yellow + primary: false = redundancy gap, not data loss. Red + primary: true = missing data for that shard.
  • One index, one cause: Pick the highest-impact index and the strongest NO decider; avoid listing every shard.
  • Cat helpers: Use GET /_cat/allocation for disk percentages per node and GET /_cat/recovery for ongoing recoveries when the decider class is unclear or recovery is in progress.

Examples

Yellow — disk watermark after node departure. Health shows yellow with unassigned replicas on logs-2025-07. Allocation explain returns primary: false, unassigned_info.reason: NODE_LEFT, but disk_threshold decider NO on es-node-2 (“above the high watermark … 90%”). Report: no data loss; root cause is disk pressure on the receiving node; remediate disk/watermark — not “node left” alone.

Red — primary with no valid copy. Health shows red on orders-2025 with one unassigned shard. Explain returns primary: true, can_allocate: no_valid_shard_copy, last_allocation_status: no_valid_shard_copy. Report: urgent; primary data missing; restore node or snapshot; mention allocate_empty_primary only as last resort with explicit data loss.

Operations

HTTP API (shorthand)elastic CLI command
GET /_cluster/healthelastic es cluster health
GET /_cluster/health?level=indiceselastic es cluster health --level indices
POST /_cluster/allocation/explainelastic es cluster allocation-explain
POST /_cluster/allocation/explain (specific shard)elastic es cluster allocation-explain --index '<index>' --shard <id> --primary true (replica: false)
GET /_cat/allocationelastic es cat allocation
GET /_cat/recoveryelastic es cat recovery
GET /_cat/shards/{index}?h=index,shard,prirep,state,unassigned.reasonelastic es cat shards --index '<index>' --h index,shard,prirep,state,unassigned.reason
POST /_cluster/reroute (last-resort empty primary — operator only)elastic es cluster reroute --commands '<json>'

来源与署名

来源:elastic/agent-skills位于skills/elasticsearch/elasticsearch-cluster-health提交baa5111

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 elastic/agent-skills 的技能

Elasticsearch Search Relevance

elastic

Improve Elasticsearch search relevance for content and catalog indices: pin or promote results with query rules (correct rule type, criteria, and rule-query wiring) and tune organic ranking with multi_match, field boosts, and analysis grounded in the index mapping. Use when search results rank poorly, a specific document must appear first for a query, or the user asks to tune full-text matching — not for ES|QL analytics, index ingest, or cluster health.

待分类592昨天更新

Elasticsearch Query Optimization

elastic

Diagnose slow Elasticsearch Query DSL searches and propose measured fixes. Use when a search is slow, profile output shows an expensive clause, exact-match filters sit in scoring context, or leading wildcards dominate latency. Ground every recommendation in search profiling — move non-scoring clauses to filter context, eliminate leading wildcards, and re-profile to confirm improvement.

待分类592昨天更新

Elasticsearch Ingest

elastic

Load CSV and JSON files into Elasticsearch indices using the bulk API and explicit mappings when field types matter. Use when batch-importing local files, converting CSV rows or JSON arrays to NDJSON bulk format, or verifying document counts and mappings after ingest — not for Logstash pipelines, Beats, custom scripts, or index-to-index reindex.

待分类592昨天更新

Elasticsearch Index Design

elastic

根据访问模式设计和审查 Elasticsearch 索引映射,涵盖字段类型、多字段和分片设置。

Data & Analytics592昨天更新

Kibana Dashboards

elastic

Create and manage Kibana Dashboards and Lens visualizations. Use when you need to define dashboards and visualizations declaratively, version control them, or automate their deployment.

待分类592昨天更新

Kibana Anomaly Detection

elastic

用于调查、解释、排查和配置 Elastic ML 异常检测作业。

Data & Analytics592昨天更新