You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NCCL内部错误:Proxy Call to rank 0连接失败问题求助

Ray集群+PyTorch分布式任务触发ncclInternalError连接失败

问题概述

搭建2个单GPU节点的Ray集群后,运行PyTorch分布式任务时,分布式进程已成功注册,使用NCCL作为后端启动2个进程,但运行后立即触发ncclInternalError: Internal check failed. Proxy Call to rank 0 failed (Connect)错误,完全阻塞任务执行。

报错详情

NCCL初始化日志

(RayExecutor pid=508760) GPU available: True (cuda), used: True (Please ignore the previous info [GPU used: False]).
(RayExecutor pid=508760) hostssh:508760:508760 [0] NCCL INFO Bootstrap : Using enp3s0:172.16.96.59<0>
(RayExecutor pid=508760) hostssh:508760:508760 [0] NCCL INFO NET/Plugin : No plugin found (libnccl-net.so), using internal implementation
(RayExecutor pid=508760) hostssh:508760:508760 [0] NCCL INFO cudaDriverVersion 11070
(RayExecutor pid=508760) NCCL version 2.14.3+cuda11.7

完整错误栈

RayTaskError(RuntimeError): [36mray::RayExecutor.execute()[39m (pid=508760, ip=172.16.96.59, repr=<ray_lightning.launchers.utils.RayExecutor object at 0x7fa16a4327d0>)
File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/ray_lightning/launchers/utils.py", line 52, in execute
return fn(*args, **kwargs)
File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/ray_lightning/launchers/ray_launcher.py", line 301, in _wrapping_function
results = function(*args, **kwargs)
File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/pytorch_lightning/trainer/trainer.py", line 811, in _fit_impl
results = self._run(model, ckpt_path=self.ckpt_path)
File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/pytorch_lightning/trainer/trainer.py", line 1172, in _run
self.__setup_profiler()
File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/pytorch_lightning/trainer/trainer.py", line 1797, in __setup_profiler
self.profiler.setup(stage=self.state.fn._setup_fn, local_rank=local_rank, log_dir=self.log_dir)
File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/pytorch_lightning/trainer/trainer.py", line 2249, in log_dir
dirpath = self.strategy.broadcast(dirpath)
File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/pytorch_lightning/strategies/ddp_spawn.py", line 215, in broadcast
torch.distributed.broadcast_object_list(obj, src, group=_group.WORLD)
File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 2084, in broadcast_object_list
broadcast(object_sizes_tensor, src=src, group=group)
File "/home/windows/miniconda3/envs/ray/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1400, in broadcast
work = default_pg.broadcast([tensor], opts)
RuntimeError: NCCL error in: /opt/conda/conda-bld/pytorch_1670525541990/work/torch/csrc/distributed/c10d/ProcessGroupNCCL.cpp:1269, internal error, NCCL version 2.14.3
ncclInternalError: Internal check failed.
Last error: Proxy Call to rank 0 failed (Connect)

环境信息

  • 部署环境:本地自建集群,无容器化
  • 依赖版本:torch=1.13.1,ray=2.3.0,ray_lightning=0.3.0
  • 系统信息:Linux Mint 20.3,Intel Core i7-11700,内存64GB
  • 单GPU代码可正常运行(批量大小16)

复现代码

from pytorch_lightning import Trainer
from torch.utils.data import DataLoader

import ray

ray.init(runtime_env={"working_dir": utils.ROOT_PATH})

dataset_params = utils.config_parse('AUTOENCODER_DATASET')
dataset = AutoEncoderDataModule(**dataset_params)
dataset.setup()

model = AutoEncoder()
autoencoder_params = utils.config_parse('AUTOENCODER_TRAIN')
print(autoencoder_params)
print(torch.cuda.device_count())
dist_env_params = utils.config_parse('DISTRIBUTED_ENV')
strategy = None
if int(dist_env_params['horovod']) == 1:
    strategy = rl.HorovodRayStrategy(use_gpu=True, num_workers=2)
elif int(dist_env_params['model_parallel']) == 1:
    strategy = rl.RayShardedStrategy(use_gpu=True, num_workers=2)
elif int(dist_env_params['data_parallel']) == 1:
    strategy = rl.RayStrategy(use_gpu=True, num_workers=2)
trainer = Trainer(**autoencoder_params,
                strategy=strategy
                )

trainer.fit(model, dataset)

PyTorch Lightning配置

[AUTOENCODER_TRAIN]
max_epochs = 100
weights_summary = full
precision = 16
gradient_clip_val = 0.0
auto_lr_find = True
auto_scale_batch_size = True
auto_select_gpus = True
check_val_every_n_epoch = 1
fast_dev_run = False
enable_progress_bar = True
detect_anomaly=True

运行命令:python run.py


排查与解决方案

1. 确认节点间网络连通性

  • 两个节点互相ping对方的内网IP(如172.16.96.59和另一节点IP),确保无丢包、延迟正常
  • 关闭节点间的防火墙,或者开放NCCL通信所需的端口范围(NCCL默认使用随机端口,可通过NCCL_PORT环境变量指定固定端口)

2. 指定NCCL通信网卡

NCCL可能自动选择了无法跨节点通信的网卡,手动指定正确的网卡:
在运行任务前设置环境变量:

export NCCL_SOCKET_IFNAME=enp3s0  # 对应日志中显示的网卡名
export NCCL_IB_DISABLE=1  # 无InfiniBand设备时禁用

3. 修正Ray集群初始化配置

  • 启动Ray head节点时,绑定到可被worker节点访问的IP:
    ray start --head --node-ip-address=你的head节点内网IP --port=6379
    
  • Worker节点启动时,指定head节点地址:
    ray start --address=head节点IP:6379 --node-ip-address=worker节点内网IP
    
  • 代码中连接到已搭建的集群,而非启动本地集群:
    ray.init(address="auto", runtime_env={"working_dir": utils.ROOT_PATH})
    

4. 检查依赖版本兼容性

  • ray_lightning 0.3.0与ray 2.3.0可能存在兼容性问题,尝试升级ray_lightning到0.4.0,或者将Ray降级到2.2.0
  • 确认NCCL版本与CUDA版本匹配,当前NCCL 2.14.3+cuda11.7与torch 1.13.1(对应CUDA11.7)兼容,若问题仍存在,可尝试更新NCCL到最新兼容版本

5. 调整PyTorch Lightning配置

  • 关闭auto_select_gpus,手动指定GPU,或者设置CUDA_VISIBLE_DEVICES环境变量确保每个进程绑定正确的GPU
  • 暂时禁用auto_lr_find和auto_scale_batch_size,先确保分布式通信正常后再启用这些自动配置

内容的提问来源于stack exchange,提问作者NavinKumarmMNK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 06:29:58