VMware环境下Elasticsearch 7.6集群节点无法互见问题排查
Hey there, let's break down why your two VMware-based Ubuntu nodes aren't forming a cluster, even though the same config works smoothly in containers. Here are the most likely culprits and how to fix them:
1. Firewall Blocking the Transport Port (9300)
Elasticsearch nodes rely on port 9300 (the default transport port) to communicate and form a cluster. Container environments usually leave this port open by default, but Ubuntu VMs often have ufw (Uncomplicated Firewall) enabled out of the box, which blocks traffic on 9300 by default.
To fix this, run these commands on both nodes:
# Allow inbound/outbound TCP traffic on port 9300 sudo ufw allow 9300/tcp # Verify the rule is active sudo ufw status
2. bootstrap.memory_lock Isn't Properly Enabled
Your config has bootstrap.memory_lock: true, but if the system hasn't been configured to let Elasticsearch lock memory, this will trigger startup warnings—and can disrupt cluster discovery. Container environments handle memory limits differently, so this issue doesn't surface there.
Resolve this with these steps:
- Edit
/etc/sysctl.confand add:
Apply the change immediately:vm.max_map_count=262144sudo sysctl -p - Edit
/etc/security/limits.confand add:elasticsearch soft memlock unlimited elasticsearch hard memlock unlimited - Restart Elasticsearch on both nodes:
sudo systemctl restart elasticsearch - Verify memory lock is working with:
Look forcurl -XGET 'http://<node-ip>:9200/_nodes/jvm?pretty'memlock_in_use: truein the output.
3. Network Connectivity Gaps Between Nodes
Even on the same virtual switch, it's worth confirming basic connectivity:
- From node-1, ping node-2's IP:
ping 192.168.30.15 - From node-2, ping node-1's IP:
ping 192.168.30.14 - Test port 9300 reachability with
nc:
If this fails, double-check VMware's virtual switch settings (promiscuous mode is rarely needed for same-subnet VMs, but it's worth ruling out).nc -zv 192.168.30.15 9300
4. Stale Cluster State from Asynchronous Startup
If you started one node first, it may have initialized a single-node cluster using the cluster.initial_master_nodes setting. When the second node launches later, it can't join because the cluster is already formed.
Fix this by resetting the cluster state:
- Stop Elasticsearch on both nodes:
sudo systemctl stop elasticsearch - Back up your data first to avoid loss:
sudo cp -r /var/lib/elasticsearch /var/lib/elasticsearch_backup - Clear cluster state data from both nodes:
sudo rm -rf /var/lib/elasticsearch/* - Start both nodes at the same time to let them initialize the cluster together.
Once you've addressed these points, check cluster health with:
curl -XGET 'http://192.168.30.14:9200/_cluster/health?pretty'
You should see status: green or yellow (depending on index replication) and number_of_nodes: 2.
内容的提问来源于stack exchange,提问作者Amir

