Elasticsearch 5.4集群3节点无法发送主节点加入请求,求排查思路
Hey there, let's work through your Elasticsearch 5.4 cluster issues—both the failed join requests to the master and the "no master found" problem are classic gotchas, but we can track them down systematically.
First, let's rule out the most common blockers that prevent nodes from reaching the master:
Verify Network & Firewall Access
Start with the basics: confirm that your data nodes can reach the master node over the cluster communication port (default 9300). Use these commands from a data node:# Test ping connectivity ping <master-node-ip> # Test port accessibility telnet <master-node-ip> 9300 # Or with curl if telnet isn't installed curl -v telnet://<master-node-ip>:9300If either test fails, you've got a network or firewall issue. Make sure 9300 is open for bidirectional traffic between all cluster nodes.
Validate elasticsearch.yml Configurations
Double-check these critical settings on all nodes:discovery.zen.ping.unicast.hosts: This must list the master node's IP/hostname and port (9300). Example:
Avoid usingdiscovery.zen.ping.unicast.hosts: ["10.0.0.5:9300", "10.0.0.6:9300"]localhosthere unless all nodes are on the same machine (which isn't typical for a 3-node cluster).network.host: Set this to an IP that's accessible to other cluster nodes (not127.0.0.1). For testing, you can use0.0.0.0, but use a specific internal IP in production.discovery.zen.minimum_master_nodes: For a 3-node cluster, this should be2(calculated as(number of master-eligible nodes / 2) + 1). This setting must match across all nodes.
If nodes can reach each other but still can't elect a master, check these areas:
Master Node Eligibility
- Ensure at least two nodes in your cluster have
node.master: trueset (the default is true, but it might have been overridden). Your master node obviously needs this enabled, and having a second eligible node prevents downtime if the primary master goes down. - Make sure every node has a unique
node.name—duplicate names will confuse the cluster's node discovery process.
- Ensure at least two nodes in your cluster have
Tweak Discovery Timeouts
If your network has slight latency, the defaultdiscovery.zen.ping_timeout: 3smight be too short. Try increasing it to 10 seconds on all nodes:discovery.zen.ping_timeout: 10sDig Into Logs for Clues
Scour the logs on both master and data nodes for specific error messages:- Search master node logs for
master not discovered yet,failed to form cluster, ornode not eligible for master—these will tell you if the master can't see other nodes or isn't being recognized as eligible. - On data nodes, look for
failed to send join request to masterorno master found—the accompanying error details (like SSL handshake failures or permission issues) will point you to the root cause.
- Search master node logs for
Check Disk Space
Elasticsearch will enter read-only mode if disk usage exceeds 90%, which can break master node functionality. Run this command on each node to verify:df -hFree up space if any node is hitting the threshold.
If you've gone through all these steps and still have issues, share redacted snippets of your elasticsearch.yml files and relevant log lines (remove any sensitive IPs or credentials)—that will help narrow things down even faster.
内容的提问来源于stack exchange,提问作者Javeed Shakeel

