You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用torchrun多GPU训练时遭遇RendezvousTimeoutError求助

PyTorch多GPU训练torchrun报错问题

问题背景

使用PyTorch的torchrun进行多GPU训练时,先后遇到RendezvousTimeoutError和RuntimeError: CUDA error: invalid device ordinal,单GPU模式可正常运行,但无法实现多GPU训练。

操作详情

创建distributed.sh脚本,内容如下:

export CUDA_VISIBLE_DEVICES=0,1
NPROC_PER_NODE=1
export MASTER_PORT=29502

# Set a smaller batch size for debugging
per_gpu_bs=4
acc_step=1
bs=$per_gpu_bs

SCRIPT_DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" && pwd )"
# Get the root directory of the project (two levels up from script)
PROJECT_ROOT="$( cd "$SCRIPT_DIR/../.." && pwd )"
export PYTHONPATH="$PROJECT_ROOT:$PYTHONPATH"

echo "PER_DEVICE_TRAIN_BATCH_SIZE="$bs


torchrun --standalone --nproc_per_node=$NPROC_PER_NODE --nnodes=2  \
    --master_port=29502 \
    --rdzv-backend=c10d \
    -m vila_u.train.train_mem_rc \
    [train_mem_rc args...]

报错1:RendezvousTimeoutError

执行上述脚本后,报错栈如下:

Traceback (most recent call last):
  File "/conda/envs/new_pipe/bin/torchrun", line 8, in <module>
    sys.exit(main())
             ^^^^^^
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 355, in wrapper
    return f(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/run.py", line 919, in main
    run(args)
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/run.py", line 910, in run
    elastic_launch(
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/launcher/api.py", line 138, in __call__
    return launch_agent(self._config, self._entrypoint, list(args))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/launcher/api.py", line 260, in launch_agent
    result = agent.run()
             ^^^^^^^^^^^
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/metrics/api.py", line 137, in wrapper
    result = f(*args, **kwargs)
             ^^^^^^^^^^^^^^^^^^
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/agent/server/api.py", line 696, in run
    result = self._invoke_run(role)
             ^^^^^^^^^^^^^^^^^^^^^^
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/agent/server/api.py", line 849, in _invoke_run
    self._initialize_workers(self._worker_group)
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/metrics/api.py", line 137, in wrapper
    result = f(*args, **kwargs)
             ^^^^^^^^^^^^^^^^^^
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/agent/server/api.py", line 668, in _initialize_workers
    self._rendezvous(worker_group)
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/metrics/api.py", line 137, in wrapper
    result = f(*args, **kwargs)
             ^^^^^^^^^^^^^^^^^^
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/agent/server/api.py", line 500, in _rendezvous
    rdzv_info = spec.rdzv_handler.next_rendezvous()
                ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py", line 1157, in next_rendezvous
    self._op_executor.run(join_op, deadline, self._get_deadline)
  File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py", line 679, in run
    raise RendezvousTimeoutError
torch.distributed.elastic.rendezvous.api.RendezvousTimeoutError

执行nvidia-smi显示服务器无运行进程:

±--------------------------------------------------------------------------------------+
| NVIDIA-SMI 535.183.01 Driver Version: 535.183.01 CUDA Version: 12.2 |
|-----------------------------------------±---------------------±---------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+======================+======================|
| 0 NVIDIA H100 80GB HBM3 Off | 00000000:1A:00.0 Off | 0 |
| N/A 31C P0 69W / 700W | 0MiB / 81559MiB | 0% Default |
| | | Disabled |
±----------------------------------------±---------------------±---------------------+
| 1 NVIDIA H100 80GB HBM3 Off | 00000000:40:00.0 Off | 0 |
| N/A 28C P0 70W / 700W | 0MiB / 81559MiB | 0% Default |
| | | Disabled |
±----------------------------------------±---------------------±---------------------+
| 2 NVIDIA H100 80GB HBM3 Off | 00000000:53:00.0 Off | 0 |
| N/A 28C P0 72W / 700W | 0MiB / 81559MiB | 0% Default |
| | | Disabled |
±----------------------------------------±---------------------±---------------------+
| 3 NVIDIA H100 80GB HBM3 Off | 00000000:66:00.0 Off | 0 |
| N/A 31C P0 70W / 700W | 0MiB / 81559MiB | 0% Default |
| | | Disabled |
±----------------------------------------±---------------------±---------------------+
| 4 NVIDIA H100 80GB HBM3 Off | 00000000:9C:00.0 Off | 0 |
| N/A 32C P0 69W / 700W | 0MiB / 81559MiB | 0% Default |
| | | Disabled |
±----------------------------------------±---------------------±---------------------+
| 5 NVIDIA H100 80GB HBM3 Off | 00000000:C0:00.0 Off | 0 |
| N/A 27C P0 69W / 700W | 0MiB / 81559MiB | 0% Default |
| | | Disabled |
±----------------------------------------±---------------------±---------------------+
| 6 NVIDIA H100 80GB HBM3 Off | 00000000:D2:00.0 Off | 0 |
| N/A 33C P0 71W / 700W | 0MiB / 81559MiB | 0% Default |
| | | Disabled |
±----------------------------------------±---------------------±---------------------+
| 7 NVIDIA H100 80GB HBM3 Off | 00000000:E4:00.0 Off | 0 |
| N/A 27C P0 72W / 700W | 0MiB / 81559MiB | 0% Default |
| | | Disabled |
±----------------------------------------±---------------------±---------------------+

±--------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=======================================================================================|
| No running processes found |
±--------------------------------------------------------------------------------------+

调试尝试及报错2:CUDA设备序数无效

  • 设置export CUDA_VISIBLE_DEVICES=0, NPROC_PER_NODE=1, nnodes=1时,脚本可正常运行,但仅单GPU训练,无法利用多GPU资源。
  • 设置CUDA_VISIBLE_DEVICES=1(或0)、NPROC_PER_NODE=2、nnodes=1时,出现RuntimeError: CUDA error: invalid device ordinal,报错栈如下:
[rank1]: Traceback (most recent call last):
[rank1]:   File "<frozen runpy>", line 198, in _run_module_as_main
[rank1]:   File "<frozen runpy>", line 88, in _run_code
[rank1]:   File "/train/train_mem_rc.py", line 19, in <module>
[rank1]:     train()
[rank1]:   File "/train/train_rc.py", line 307, in train
[rank1]:     trainer.train(resume_from_checkpoint=resume_from_checkpoint)
[rank1]:   File "/conda/envs/new_pipe/lib/python3.12/site-packages/transformers/trainer.py", line 2247, in train
[rank1]:     return inner_training_loop(
[rank1]:            ^^^^^^^^^^^^^^^^^^^^
[rank1]:   File "/conda/envs/new_pipe/lib/python3.12/site-packages/transformers/trainer.py", line 2376, in _inner_training_loop
[rank1]:     model, self.optimizer = self.accelerator.prepare(self.model, self.optimizer)
[rank1]:                             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]:   File "/conda/envs/new_pipe/lib/python3.12/site-packages/accelerate/accelerator.py", line 1318, in prepare
[rank1]:     result = self._prepare_deepspeed(*args)
[rank1]:              ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]:   File "/conda/envs/new_pipe/lib/python3.12/site-packages/accelerate/accelerator.py", line 1815, in _prepare_deepspeed
[rank1]:     engine, optimizer, _, lr_scheduler = deepspeed.initialize(**kwargs)
[rank1]:                                          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]:   File "/conda/envs/new_pipe/lib/python3.12/site-packages/deepspeed/__init__.py", line 193, in initialize
[rank1]:     engine = DeepSpeedEngine(args=args,
[rank1]:              ^^^^^^^^^^^^^^^^^^^^^^^^^^
[rank1]:   File "/conda/envs/new_pipe/lib/python3.12/site-packages/deepspeed/runtime/engine.py", line 264, in __init__
[rank1]:     self._set_distributed_vars(args)
[rank1]:   File "/conda/envs/new_pipe/lib/python3.12/site-packages/deepspeed/runtime/engine.py", line 1121, in _set_distributed_vars
[rank1]:     get_accelerator().set_device(device_rank)
[rank1]:   File "/conda/envs/new_pipe/lib/python3.12/site-packages/deepspeed/accelerator/cuda_accelerator.py", line 67, in set_device
[rank1]:     torch.cuda.set_device(device_index)
[rank1]:   File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/cuda/__init__.py", line 478, in set_device
[rank1]:     torch._C._cuda_setDevice(device)
[rank1]: RuntimeError: CUDA error: invalid device ordinal
[rank1]: CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
[rank1]: For debugging consider passing CUDA_LAUNCH_BLOCKING=1
[rank1]: Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.

疑问

是操作参数设置有误,还是服务器性能问题?如何调整才能实现多GPU正常训练?


内容的提问来源于stack exchange,提问作者Sofia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 16:47:02