Elasticsearch集群变红求助:主分片分配失败(v5.0生产双节点)
Alright, let's tackle this cluster red status issue you're facing. From the GET /_cluster/allocation/explain output you shared, the core problem is straightforward:
failed to create shard, failure IOException[failed to obtain in-memory shard lock]; nested: NotSerializableExceptionWrapper[shard_lock_obtain_failed_exception: [twitter_tracker][0]: obtaining shard lock timed out after 5000ms]
Your primary shard 0 for twitter_tracker can't acquire the in-memory shard lock within the default 5-second timeout. The good news is that node elasticsearch-24-384-node-1 has a valid, available copy of this shard—we can leverage that to restore the cluster.
Step 1: Force Allocate the Stale Primary Shard to the Node with Valid Data
Since node 1 already has the shard data, we can manually reroute the primary shard to it. First, make sure cluster allocation is enabled (it might have been restricted temporarily):
PUT /_cluster/settings { "transient": { "cluster.routing.allocation.enable": "all" } }
Then run the reroute command to assign the stale primary shard to node 1. The accept_data_loss flag is necessary here because we're using a potentially non-fresh copy, but since it's the only available valid copy, this is safe:
POST /_cluster/reroute { "commands": [ { "allocate_stale_primary": { "index": "twitter_tracker", "shard": 0, "node": "elasticsearch-24-384-node-1", "accept_data_loss": true } } ] }
Step 2: Address the Root Cause of the Shard Lock Timeout
To prevent this from happening again, check these key areas:
- JVM Heap Memory: If node 1's heap is overloaded, it can slow down lock operations. Use
GET /_nodes/jvmto check heap usage—aim to keep it below 70-80% of the allocated heap. - Disk IO Performance: High disk latency can delay lock acquisition. Use system tools like
iostatortopto monitor disk read/write speeds on your nodes. - Adjust Lock Timeout: If your cluster regularly deals with heavy loads, increase the shard lock wait time for the affected index:
PUT /twitter_tracker/_settings { "index.shard.lock.wait_time": "10s" }
Step 3: Verify Cluster Recovery
After running the above commands, confirm the cluster is back to green:
GET /_cluster/health GET /twitter_tracker/_shard/0/_status
You should see the cluster status as green and the shard status as STARTED.
内容的提问来源于stack exchange,提问作者Tamizharasan

