You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

分布式TensorFlow运行报错UnavailableError: Endpoint read fail求助

Troubleshooting "UnavailableError: Endpoint read failed" in Distributed TensorFlow

Hey there, let's break down the issue you're facing with your distributed TensorFlow setup. I've gone through your code and logs, and here are some common causes and actionable fixes to try out:

1. Confirm Both Server Processes Are Running Properly

Looking at your server logs, I only see output for server #0 (port 2222) plus a "Server already started" message—this suggests you might have accidentally launched two instances of task 0 instead of task 0 and task 1.

  • Double-check your terminal commands:
    • Terminal 1: Run python your_server_script.py 0 (should start server #0 on port 2222)
    • Terminal 2: Run python your_server_script.py 1 (should start server #1 on port 2223)
  • Ensure both terminals show their respective "Starting server #X" message and gRPC startup logs without errors.
  • Use ps aux | grep python in a third terminal to confirm both Python processes are active in the background.

2. Test Port Connectivity

Even for localhost, loopback port issues can block gRPC communication. Let's verify the ports are reachable:

  • Run these commands in a new terminal:
    nc -zv localhost 2222
    nc -zv localhost 2223
    
  • If you see "Connection succeeded" for both, the ports are open. If not:
    • Use lsof -i :2222 and lsof -i :2223 to check if another process is using those ports.
    • Kill any conflicting processes and restart your TensorFlow servers.

3. Simplify Your Test Code to Isolate the Issue

Your current calculation code specifies device placement for tasks 0 and 1. Let's rule out device assignment as the culprit with a minimal working example:

import tensorflow as tf

# Reuse the same cluster spec as your server code
cluster = tf.train.ClusterSpec({"local": ["localhost:2222", "localhost:2223"]})

# Try connecting to server #0 first
with tf.Session("grpc://localhost:2222") as sess:
    simple_sum = tf.constant(10) + tf.constant(20)
    result = sess.run(simple_sum)
    print("Simple sum result:", result)

If this works, the core distributed setup is functional, and the issue likely lies in your device placement logic. If it still throws the same error, focus on fixing server communication first.

4. Check TensorFlow and gRPC Version Compatibility

TensorFlow 1.x (which you're using) has strict version requirements for gRPC. Mismatched versions often cause communication failures:

  • Check your current versions:
    pip list | grep -E "tensorflow|grpcio"
    
  • For TensorFlow 1.8-1.10, compatible gRPC versions are typically 1.10.x-1.12.x. If your versions are way off, install matching versions:
    pip install tensorflow==1.10 grpcio==1.12.1 --force-reinstall
    

5. Disable Local Firewall/Security Software

Some local firewalls or security tools block loopback (localhost) traffic by default. Temporarily disable your firewall to test:

  • On Ubuntu/Debian: sudo ufw disable
  • On CentOS/RHEL: sudo systemctl stop firewalld
  • After testing, re-enable it with sudo ufw enable or sudo systemctl start firewalld.

6. Verify Session Target Correctness

Your code connects to grpc://localhost:2223 (server #1). Make sure server #1 is actually running and listening on that port. Try switching the session target to grpc://localhost:2222 (server #0) to see if that works—this can tell you if one of the servers is misconfigured.

Give these steps a try, and let me know which one resolves your issue!

内容的提问来源于stack exchange,提问作者Jekyll SONG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:18:43