分布式TensorFlow运行报错UnavailableError: Endpoint read fail求助
Hey there, let's break down the issue you're facing with your distributed TensorFlow setup. I've gone through your code and logs, and here are some common causes and actionable fixes to try out:
1. Confirm Both Server Processes Are Running Properly
Looking at your server logs, I only see output for server #0 (port 2222) plus a "Server already started" message—this suggests you might have accidentally launched two instances of task 0 instead of task 0 and task 1.
- Double-check your terminal commands:
- Terminal 1: Run
python your_server_script.py 0(should start server #0 on port 2222) - Terminal 2: Run
python your_server_script.py 1(should start server #1 on port 2223)
- Terminal 1: Run
- Ensure both terminals show their respective "Starting server #X" message and gRPC startup logs without errors.
- Use
ps aux | grep pythonin a third terminal to confirm both Python processes are active in the background.
2. Test Port Connectivity
Even for localhost, loopback port issues can block gRPC communication. Let's verify the ports are reachable:
- Run these commands in a new terminal:
nc -zv localhost 2222 nc -zv localhost 2223 - If you see "Connection succeeded" for both, the ports are open. If not:
- Use
lsof -i :2222andlsof -i :2223to check if another process is using those ports. - Kill any conflicting processes and restart your TensorFlow servers.
- Use
3. Simplify Your Test Code to Isolate the Issue
Your current calculation code specifies device placement for tasks 0 and 1. Let's rule out device assignment as the culprit with a minimal working example:
import tensorflow as tf # Reuse the same cluster spec as your server code cluster = tf.train.ClusterSpec({"local": ["localhost:2222", "localhost:2223"]}) # Try connecting to server #0 first with tf.Session("grpc://localhost:2222") as sess: simple_sum = tf.constant(10) + tf.constant(20) result = sess.run(simple_sum) print("Simple sum result:", result)
If this works, the core distributed setup is functional, and the issue likely lies in your device placement logic. If it still throws the same error, focus on fixing server communication first.
4. Check TensorFlow and gRPC Version Compatibility
TensorFlow 1.x (which you're using) has strict version requirements for gRPC. Mismatched versions often cause communication failures:
- Check your current versions:
pip list | grep -E "tensorflow|grpcio" - For TensorFlow 1.8-1.10, compatible gRPC versions are typically 1.10.x-1.12.x. If your versions are way off, install matching versions:
pip install tensorflow==1.10 grpcio==1.12.1 --force-reinstall
5. Disable Local Firewall/Security Software
Some local firewalls or security tools block loopback (localhost) traffic by default. Temporarily disable your firewall to test:
- On Ubuntu/Debian:
sudo ufw disable - On CentOS/RHEL:
sudo systemctl stop firewalld - After testing, re-enable it with
sudo ufw enableorsudo systemctl start firewalld.
6. Verify Session Target Correctness
Your code connects to grpc://localhost:2223 (server #1). Make sure server #1 is actually running and listening on that port. Try switching the session target to grpc://localhost:2222 (server #0) to see if that works—this can tell you if one of the servers is misconfigured.
Give these steps a try, and let me know which one resolves your issue!
内容的提问来源于stack exchange,提问作者Jekyll SONG

