Docker容器中PyTorch torchrun非本地IP连接失败求助
问题
在Docker容器中练习PyTorch多节点DDP时,执行以下torchrun命令(--rdzv-endpoint设为localhost:10000)程序运行正常:
torchrun \ --nnodes=1 \ --node_rank=0 \ --nproc_per_node=gpu \ --rdzv_id=123 \ --rdzv-backend=c10d \ --rdzv-endpoint=localhost:10000 \ test_code.py
但将--rdzv-endpoint改为机器的192.168.9.225:10000后,程序卡住并抛出RendezvousConnectionError错误,具体报错信息如下:
master_addr is only used for static rdzv_backend and when rdzv_endpoint is not specified. WARNING:torch.distributed.run: ***************************************** Setting OMP_NUM_THREADS environment variable for each process to be 1 in default, to avoid your system being overloaded, please further tune the variable for optimal performance in your application as needed. ***************************************** [E socket.cpp:860] [c10d] The client socket has timed out after 60s while trying to connect to (192.168.9.225, 10000). Traceback (most recent call last): File "/opt/conda/lib/python3.10/site-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py", line 155, in _create_tcp_store store = TCPStore( TimeoutError: The client socket has timed out after 60s while trying to connect to (192.168.9.225, 10000). The above exception was the direct cause of the following exception: Traceback (most recent call last): File "/opt/conda/bin/torchrun", line 33, in <module> sys.exit(load_entry_point('torch==2.0.1', 'console_scripts', 'torchrun')()) File "/opt/conda/lib/python3.10/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 346, in wrapper return f(*args, **kwargs) File "/opt/conda/lib/python3.10/site-packages/torch/distributed/run.py", line 794, in main run(args) File "/opt/conda/lib/python3.10/site-packages/torch/distributed/run.py", line 785, in run elastic_launch( File "/opt/conda/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 134, in __call__ return launch_agent(self._config, self._entrypoint, list(args)) File "/opt/conda/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 223, in launch_agent rdzv_handler=rdzv_registry.get_rendezvous_handler(rdzv_parameters), File "/opt/conda/lib/python3.10/site-packages/torch/distributed/elastic/rendezvous/registry.py", line 65, in get_rendezvous_handler return handler_registry.create_handler(params) File "/opt/conda/lib/python3.10/site-packages/torch/distributed/elastic/rendezvous/api.py", line 257, in create_handler handler = creator(params) File "/opt/conda/lib/python3.10/site-packages/torch/distributed/elastic/rendezvous/registry.py", line 36, in _create_c10d_handler backend, store = create_backend(params) File "/opt/conda/lib/python3.10/site-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py", line 250, in create_backend store = _create_tcp_store(params) File "/opt/conda/lib/python3.10/site-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py", line 175, in _create_tcp_store raise RendezvousConnectionError( torch.distributed.elastic.rendezvous.api.RendezvousConnectionError: The connection to the C10d store has failed. See inner exception for details.
Docker容器创建命令如下:
docker run -it --gpus=all --ipc=host --network=host --cap-add=NET_ADMIN --name=pytorch-2.0-examples -v=pytorch-2.0-examples:/pytorch-2.0-examples pytorch/pytorch /bin/bash
已确认ping测试正常,且测试时已关闭防火墙,需解决PyTorch torchrun使用非127.0.0.1的IP地址正常运行的问题。
解决方案
- 显式指定master_addr参数:执行torchrun时添加
--master_addr=192.168.9.225,明确让进程绑定到目标网卡IP,避免默认只监听localhost。调整后的命令:
torchrun \ --nnodes=1 \ --node_rank=0 \ --nproc_per_node=gpu \ --rdzv_id=123 \ --rdzv-backend=c10d \ --rdzv-endpoint=192.168.9.225:10000 \ --master_addr=192.168.9.225 \ test_code.py
检查容器内端口监听状态:进入容器后执行
netstat -tulpn,确认端口10000的监听地址是0.0.0.0或192.168.9.225,而非仅127.0.0.1。若进程只绑定localhost,即使容器用host网络,也无法通过网卡IP访问。自定义TCPStore初始化逻辑:如果代码中手动初始化DDP,显式设置TCPStore的监听地址为
192.168.9.225,确保Store监听在目标IP上。示例代码片段:
import torch.distributed as dist from torch.distributed import TCPStore # 主节点初始化Store store = TCPStore("192.168.9.225", 10000, is_master=True, timeout=30) dist.init_process_group(backend="nccl", store=store, rank=0, world_size=4)
- 确认容器网络转发设置:在容器内执行
sysctl net.ipv4.ip_forward=1开启IP转发,同时检查宿主机对应网卡是否处于正常启用状态,避免网络包被拦截。
内容的提问来源于stack exchange,提问作者GeSol
相关产品推荐
相关产品推荐

