Docker Swarm集群升级ES镜像后应用部署连接失败求助
Let's dive into why your second container can't reach Elasticsearch after upgrading from 2.4 to 5.6 in your Docker Swarm cluster. That "None of the configured nodes are available" error is almost always tied to version-specific configuration changes or Swarm network/startup timing issues. Here's how to troubleshoot and fix it:
1. Address Elasticsearch 5.x Configuration Changes
ES 5.6 introduced critical changes that break compatibility with 2.4's defaults—these are the most likely culprits:
Network Binding: By default, ES 5.x only binds to
localhost, which means other containers on the Swarm network can't reach it. Add this environment variable to your ES service:environment: - network.host=0.0.0.0This allows ES to listen on all network interfaces in the container.
Node Discovery: ES 5.x uses
discovery.zen.ping.unicast.hostsfor node discovery (replacing older mechanisms). For a single-node cluster, set this to your ES service name:environment: - discovery.zen.ping.unicast.hosts=nga_es - discovery.zen.minimum_master_nodes=1 # Critical for preventing split-brain; use (n/2)+1 for multi-node clustersCluster Name Match: Double-check that your ES service's
cluster.namematches what your app is using. Add this to confirm:environment: - cluster.name=elasticsearch
2. Fix Docker Swarm Startup Timing & Network
ES 5.6 takes longer to initialize than 2.4, so your app container might be trying to connect before ES is ready. Plus, you need to ensure both services are on the same Swarm network:
Add Health Checks to ES: Force Swarm to wait until ES is fully ready before starting your app. Add this to your ES service:
healthcheck: test: ["CMD-SHELL", "curl -s http://localhost:9200/_cluster/health | grep -q '\"status\":\"green\"'"] interval: 10s timeout: 10s retries: 5Wait for Healthy ES in Your App: Update your app service to depend on the healthy state of the ES service (note: this works in Docker Compose v3.7+ for Swarm):
depends_on: nga_es: condition: service_healthyShared Overlay Network: Ensure both your ES and app services are attached to the same Swarm overlay network. Define it in your
docker-compose.yml:networks: swarm_internal: driver: overlayThen add
networks: [swarm_internal]to both services.
3. Verify Basic Connectivity
If the above changes don't work, run these checks to narrow down the issue:
- Check ES Service Logs: Run
docker service logs nga_esto see if ES is failing to start (common issues: memory limits, file permissions). AddES_JAVA_OPTS=-Xms512m -Xmx512mto your ES environment if you hit memory errors. - Test DNS Resolution: Exec into your app container and run
ping nga_es—if it fails, your network setup is broken. - Test Port Access: Run
telnet nga_es 9300from your app container. If it can't connect, confirm ES is listening on 9300 (checknetwork.hostagain) and that no Swarm network policies are blocking traffic.
Example Updated docker-compose.yml Snippet
Here's how your ES and app services should look with these fixes:
version: "3.8" services: nga_es: image: elasticsearch:5.6 environment: - cluster.name=elasticsearch - discovery.zen.ping.unicast.hosts=nga_es - discovery.zen.minimum_master_nodes=1 - network.host=0.0.0.0 - ES_JAVA_OPTS=-Xms512m -Xmx512m volumes: - es_data:/usr/share/elasticsearch/data networks: - swarm_internal healthcheck: test: ["CMD-SHELL", "curl -s http://localhost:9200/_cluster/health | grep -q '\"status\":\"green\"'"] interval: 10s timeout: 10s retries: 5 deploy: replicas: 1 restart_policy: condition: on-failure your_app: image: your-app-image:latest environment: - ES_HOST=nga_es - ES_PORT=9300 - ES_CLUSTER_NAME=elasticsearch networks: - swarm_internal depends_on: nga_es: condition: service_healthy deploy: restart_policy: condition: on-failure networks: swarm_internal: driver: overlay volumes: es_data: driver: local
内容的提问来源于stack exchange,提问作者thechmodmaster

