NCCL内部错误:Proxy Call to rank 0连接失败问题求助
Ray集群+PyTorch分布式任务触发ncclInternalError连接失败
问题概述
搭建2个单GPU节点的Ray集群后,运行PyTorch分布式任务时,分布式进程已成功注册,使用NCCL作为后端启动2个进程,但运行后立即触发ncclInternalError: Internal check failed. Proxy Call to rank 0 failed (Connect)错误,完全阻塞任务执行。
报错详情
NCCL初始化日志
(RayExecutor pid=508760) GPU available: True (cuda), used: True (Please ignore the previous info [GPU used: False]). (RayExecutor pid=508760) hostssh:508760:508760 [0] NCCL INFO Bootstrap : Using enp3s0:172.16.96.59<0> (RayExecutor pid=508760) hostssh:508760:508760 [0] NCCL INFO NET/Plugin : No plugin found (libnccl-net.so), using internal implementation (RayExecutor pid=508760) hostssh:508760:508760 [0] NCCL INFO cudaDriverVersion 11070 (RayExecutor pid=508760) NCCL version 2.14.3+cuda11.7
完整错误栈
RayTaskError(RuntimeError): [36mray::RayExecutor.execute()[39m (pid=508760, ip=172.16.96.59, repr=<ray_lightning.launchers.utils.RayExecutor object at 0x7fa16a4327d0>) File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/ray_lightning/launchers/utils.py", line 52, in execute return fn(*args, **kwargs) File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/ray_lightning/launchers/ray_launcher.py", line 301, in _wrapping_function results = function(*args, **kwargs) File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/pytorch_lightning/trainer/trainer.py", line 811, in _fit_impl results = self._run(model, ckpt_path=self.ckpt_path) File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/pytorch_lightning/trainer/trainer.py", line 1172, in _run self.__setup_profiler() File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/pytorch_lightning/trainer/trainer.py", line 1797, in __setup_profiler self.profiler.setup(stage=self.state.fn._setup_fn, local_rank=local_rank, log_dir=self.log_dir) File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/pytorch_lightning/trainer/trainer.py", line 2249, in log_dir dirpath = self.strategy.broadcast(dirpath) File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/pytorch_lightning/strategies/ddp_spawn.py", line 215, in broadcast torch.distributed.broadcast_object_list(obj, src, group=_group.WORLD) File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 2084, in broadcast_object_list broadcast(object_sizes_tensor, src=src, group=group) File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1400, in broadcast work = default_pg.broadcast([tensor], opts) RuntimeError: NCCL error in: /opt/conda/conda-bld/pytorch_1670525541990/work/torch/csrc/distributed/c10d/ProcessGroupNCCL.cpp:1269, internal error, NCCL version 2.14.3 ncclInternalError: Internal check failed. Last error: Proxy Call to rank 0 failed (Connect)
环境信息
- 部署环境:本地自建集群,无容器化
- 依赖版本:torch=1.13.1,ray=2.3.0,ray_lightning=0.3.0
- 系统信息:Linux Mint 20.3,Intel Core i7-11700,内存64GB
- 单GPU代码可正常运行(批量大小16)
复现代码
from pytorch_lightning import Trainer from torch.utils.data import DataLoader import ray ray.init(runtime_env={"working_dir": utils.ROOT_PATH}) dataset_params = utils.config_parse('AUTOENCODER_DATASET') dataset = AutoEncoderDataModule(**dataset_params) dataset.setup() model = AutoEncoder() autoencoder_params = utils.config_parse('AUTOENCODER_TRAIN') print(autoencoder_params) print(torch.cuda.device_count()) dist_env_params = utils.config_parse('DISTRIBUTED_ENV') strategy = None if int(dist_env_params['horovod']) == 1: strategy = rl.HorovodRayStrategy(use_gpu=True, num_workers=2) elif int(dist_env_params['model_parallel']) == 1: strategy = rl.RayShardedStrategy(use_gpu=True, num_workers=2) elif int(dist_env_params['data_parallel']) == 1: strategy = rl.RayStrategy(use_gpu=True, num_workers=2) trainer = Trainer(**autoencoder_params, strategy=strategy ) trainer.fit(model, dataset)
PyTorch Lightning配置
[AUTOENCODER_TRAIN] max_epochs = 100 weights_summary = full precision = 16 gradient_clip_val = 0.0 auto_lr_find = True auto_scale_batch_size = True auto_select_gpus = True check_val_every_n_epoch = 1 fast_dev_run = False enable_progress_bar = True detect_anomaly=True
运行命令:python run.py
排查与解决方案
1. 确认节点间网络连通性
- 两个节点互相ping对方的内网IP(如172.16.96.59和另一节点IP),确保无丢包、延迟正常
- 关闭节点间的防火墙,或者开放NCCL通信所需的端口范围(NCCL默认使用随机端口,可通过
NCCL_PORT环境变量指定固定端口)
2. 指定NCCL通信网卡
NCCL可能自动选择了无法跨节点通信的网卡,手动指定正确的网卡:
在运行任务前设置环境变量:
export NCCL_SOCKET_IFNAME=enp3s0 # 对应日志中显示的网卡名 export NCCL_IB_DISABLE=1 # 无InfiniBand设备时禁用
3. 修正Ray集群初始化配置
- 启动Ray head节点时,绑定到可被worker节点访问的IP:
ray start --head --node-ip-address=你的head节点内网IP --port=6379 - Worker节点启动时,指定head节点地址:
ray start --address=head节点IP:6379 --node-ip-address=worker节点内网IP - 代码中连接到已搭建的集群,而非启动本地集群:
ray.init(address="auto", runtime_env={"working_dir": utils.ROOT_PATH})
4. 检查依赖版本兼容性
- ray_lightning 0.3.0与ray 2.3.0可能存在兼容性问题,尝试升级ray_lightning到0.4.0,或者将Ray降级到2.2.0
- 确认NCCL版本与CUDA版本匹配,当前NCCL 2.14.3+cuda11.7与torch 1.13.1(对应CUDA11.7)兼容,若问题仍存在,可尝试更新NCCL到最新兼容版本
5. 调整PyTorch Lightning配置
- 关闭
auto_select_gpus,手动指定GPU,或者设置CUDA_VISIBLE_DEVICES环境变量确保每个进程绑定正确的GPU - 暂时禁用
auto_lr_find和auto_scale_batch_size,先确保分布式通信正常后再启用这些自动配置
内容的提问来源于stack exchange,提问作者NavinKumarmMNK
相关产品推荐
相关产品推荐

