如何解决两台GCP VM的Docker容器中PyTorch分布式训练连接失败问题
跨GCP VM容器分布式训练NCCL连接失败解决方法
问题背景
两台GCP VM,每台运行Docker容器,容器启动命令:
docker run --gpus all -it --rm --entrypoint /bin/bash -p 8000:8000 -p 7860:7860 -p 29500:29500 lf
尝试用Llama Factory进行分布式训练,Rank1容器执行:
FORCE_TORCHRUN=1 NNODES=2 RANK=1 MASTER_ADDR=34.138.7.129 MASTER_PORT=29500 llamafactory-cli train examples/train_lora/llama3_lora_sft_ds3.yaml
Rank0容器执行:
FORCE_TORCHRUN=1 NNODES=2 RANK=0 MASTER_ADDR=34.138.7.129 MASTER_PORT=29500 llamafactory-cli train examples/train_lora/llama3_lora_sft_ds3.yaml
出现错误:
[rank1]: torch.distributed.DistBackendError: NCCL error in: ../torch/csrc/distributed/c10d/ProcessGroupNCCL.cpp:1970, unhandled system error (run with NCCL_DEBUG=INFO for details), NCCL version 2.20.5 [rank1]: ncclSystemError: System call (e.g. socket, malloc) or external library call failed or device error. [rank1]: Last error: [rank1]: socketStartConnect: Connect to 172.17.0.2<49113> failed : Software caused connection abort E0924 21:26:39.866000 140711615779968 torch/distributed/elastic/multiprocessing/api.py:826] failed (exitcode: 1) local_rank: 0 (pid: 484) of binary: /usr/bin/python3.10 Traceback (most recent call last): File "/usr/local/bin/torchrun", line 8, in <module> sys.exit(main()) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 347, in wrapper return f(*args, **kwargs) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/run.py", line 879, in main run(args) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/run.py", line 870, in run elastic_launch( File "/usr/local/lib/python3.10/dist-packages/torch/distributed/launcher/api.py", line 132, in __call__ return launch_agent(self._config, self._entrypoint, list(args)) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/launcher/api.py", line 263, in launch_agent raise ChildFailedError( torch.distributed.elastic.multiprocessing.errors.ChildFailedError: ============================================================ /workspace/LLaMA-Factory/src/llamafactory/launcher.py FAILED ------------------------------------------------------------ Failures: <NO_OTHER_FAILURES> ------------------------------------------------------------ Root Cause (first observed failure): [0]: time : 2024-09-24_21:26:39 host : 71af1f49abe3 rank : 1 (local_rank: 0) exitcode : 1 (pid: 484) error_file: <N/A> traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html ============================================================
问题核心:PyTorch/NCCL默认使用容器内部IP(172.17.0.2)进行通信,而非GCP VM的公网/内部IP,导致跨VM容器无法连通。
解决方案
1. 使用Docker Host网络模式
让容器直接使用VM的网络栈,避免IP映射问题。修改容器启动命令:
docker run --gpus all -it --rm --entrypoint /bin/bash --net=host lf
- 无需再指定
-p端口映射,容器直接复用VM的端口。 - 此方法最直接,能彻底解决容器IP与VMIP不一致的问题。
2. 指定NCCL绑定的网卡/IP
如果不想使用host模式,可通过环境变量强制NCCL使用VM的物理网卡:
- 先确认VM的网卡名称(GCP默认是
eth0,可通过ifconfig或ip addr查看) - 在训练命令中添加
NCCL_SOCKET_IFNAME=eth0参数:
Rank0容器命令:
FORCE_TORCHRUN=1 NNODES=2 RANK=0 MASTER_ADDR=34.138.7.129 MASTER_PORT=29500 NCCL_SOCKET_IFNAME=eth0 llamafactory-cli train examples/train_lora/llama3_lora_sft_ds3.yaml
Rank1容器命令:
FORCE_TORCHRUN=1 NNODES=2 RANK=1 MASTER_ADDR=34.138.7.129 MASTER_PORT=29500 NCCL_SOCKET_IFNAME=eth0 llamafactory-cli train examples/train_lora/llama3_lora_sft_ds3.yaml
若仍有问题,可禁用IB(InfiniBand)进一步强制使用TCP:
FORCE_TORCHRUN=1 NNODES=2 RANK=0 MASTER_ADDR=34.138.7.129 MASTER_PORT=29500 NCCL_IB_DISABLE=1 NCCL_SOCKET_IFNAME=eth0 llamafactory-cli train examples/train_lora/llama3_lora_sft_ds3.yaml
3. 配置GCP VM防火墙规则
确保两台VM之间的通信端口开放:
- 在GCP控制台创建防火墙规则,允许两台VM的公网IP/内部IP之间的TCP/UDP流量,端口范围包含
29500以及NCCL可能用到的动态端口(或直接允许所有内部流量,更便捷)。
4. 优先使用GCP内部IP
公网IP存在延迟和防火墙限制,建议改用VM的内部私有IP作为MASTER_ADDR,流量在GCP内部网络传输,稳定性和速度更优。
内容的提问来源于stack exchange,提问作者BAE
相关产品推荐
相关产品推荐

