使用torchrun多GPU训练时遭遇RendezvousTimeoutError求助
PyTorch多GPU训练torchrun报错问题
问题背景
使用PyTorch的torchrun进行多GPU训练时,先后遇到RendezvousTimeoutError和RuntimeError: CUDA error: invalid device ordinal,单GPU模式可正常运行,但无法实现多GPU训练。
操作详情
创建distributed.sh脚本,内容如下:
export CUDA_VISIBLE_DEVICES=0,1 NPROC_PER_NODE=1 export MASTER_PORT=29502 # Set a smaller batch size for debugging per_gpu_bs=4 acc_step=1 bs=$per_gpu_bs SCRIPT_DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" && pwd )" # Get the root directory of the project (two levels up from script) PROJECT_ROOT="$( cd "$SCRIPT_DIR/../.." && pwd )" export PYTHONPATH="$PROJECT_ROOT:$PYTHONPATH" echo "PER_DEVICE_TRAIN_BATCH_SIZE="$bs torchrun --standalone --nproc_per_node=$NPROC_PER_NODE --nnodes=2 \ --master_port=29502 \ --rdzv-backend=c10d \ -m vila_u.train.train_mem_rc \ [train_mem_rc args...]
报错1:RendezvousTimeoutError
执行上述脚本后,报错栈如下:
Traceback (most recent call last): File "/conda/envs/new_pipe/bin/torchrun", line 8, in <module> sys.exit(main()) ^^^^^^ File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 355, in wrapper return f(*args, **kwargs) ^^^^^^^^^^^^^^^^^^ File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/run.py", line 919, in main run(args) File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/run.py", line 910, in run elastic_launch( File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/launcher/api.py", line 138, in __call__ return launch_agent(self._config, self._entrypoint, list(args)) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/launcher/api.py", line 260, in launch_agent result = agent.run() ^^^^^^^^^^^ File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/metrics/api.py", line 137, in wrapper result = f(*args, **kwargs) ^^^^^^^^^^^^^^^^^^ File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/agent/server/api.py", line 696, in run result = self._invoke_run(role) ^^^^^^^^^^^^^^^^^^^^^^ File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/agent/server/api.py", line 849, in _invoke_run self._initialize_workers(self._worker_group) File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/metrics/api.py", line 137, in wrapper result = f(*args, **kwargs) ^^^^^^^^^^^^^^^^^^ File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/agent/server/api.py", line 668, in _initialize_workers self._rendezvous(worker_group) File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/metrics/api.py", line 137, in wrapper result = f(*args, **kwargs) ^^^^^^^^^^^^^^^^^^ File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/agent/server/api.py", line 500, in _rendezvous rdzv_info = spec.rdzv_handler.next_rendezvous() ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py", line 1157, in next_rendezvous self._op_executor.run(join_op, deadline, self._get_deadline) File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py", line 679, in run raise RendezvousTimeoutError torch.distributed.elastic.rendezvous.api.RendezvousTimeoutError
执行nvidia-smi显示服务器无运行进程:
±--------------------------------------------------------------------------------------+ | NVIDIA-SMI 535.183.01 Driver Version: 535.183.01 CUDA Version: 12.2 | |-----------------------------------------±---------------------±---------------------+ | GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+======================+======================| | 0 NVIDIA H100 80GB HBM3 Off | 00000000:1A:00.0 Off | 0 | | N/A 31C P0 69W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | ±----------------------------------------±---------------------±---------------------+ | 1 NVIDIA H100 80GB HBM3 Off | 00000000:40:00.0 Off | 0 | | N/A 28C P0 70W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | ±----------------------------------------±---------------------±---------------------+ | 2 NVIDIA H100 80GB HBM3 Off | 00000000:53:00.0 Off | 0 | | N/A 28C P0 72W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | ±----------------------------------------±---------------------±---------------------+ | 3 NVIDIA H100 80GB HBM3 Off | 00000000:66:00.0 Off | 0 | | N/A 31C P0 70W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | ±----------------------------------------±---------------------±---------------------+ | 4 NVIDIA H100 80GB HBM3 Off | 00000000:9C:00.0 Off | 0 | | N/A 32C P0 69W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | ±----------------------------------------±---------------------±---------------------+ | 5 NVIDIA H100 80GB HBM3 Off | 00000000:C0:00.0 Off | 0 | | N/A 27C P0 69W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | ±----------------------------------------±---------------------±---------------------+ | 6 NVIDIA H100 80GB HBM3 Off | 00000000:D2:00.0 Off | 0 | | N/A 33C P0 71W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | ±----------------------------------------±---------------------±---------------------+ | 7 NVIDIA H100 80GB HBM3 Off | 00000000:E4:00.0 Off | 0 | | N/A 27C P0 72W / 700W | 0MiB / 81559MiB | 0% Default | | | | Disabled | ±----------------------------------------±---------------------±---------------------+ ±--------------------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=======================================================================================| | No running processes found | ±--------------------------------------------------------------------------------------+
调试尝试及报错2:CUDA设备序数无效
- 设置
export CUDA_VISIBLE_DEVICES=0, NPROC_PER_NODE=1, nnodes=1时,脚本可正常运行,但仅单GPU训练,无法利用多GPU资源。 - 设置
CUDA_VISIBLE_DEVICES=1(或0)、NPROC_PER_NODE=2、nnodes=1时,出现RuntimeError: CUDA error: invalid device ordinal,报错栈如下:
[rank1]: Traceback (most recent call last): [rank1]: File "<frozen runpy>", line 198, in _run_module_as_main [rank1]: File "<frozen runpy>", line 88, in _run_code [rank1]: File "/train/train_mem_rc.py", line 19, in <module> [rank1]: train() [rank1]: File "/train/train_rc.py", line 307, in train [rank1]: trainer.train(resume_from_checkpoint=resume_from_checkpoint) [rank1]: File "/conda/envs/new_pipe/lib/python3.12/site-packages/transformers/trainer.py", line 2247, in train [rank1]: return inner_training_loop( [rank1]: ^^^^^^^^^^^^^^^^^^^^ [rank1]: File "/conda/envs/new_pipe/lib/python3.12/site-packages/transformers/trainer.py", line 2376, in _inner_training_loop [rank1]: model, self.optimizer = self.accelerator.prepare(self.model, self.optimizer) [rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [rank1]: File "/conda/envs/new_pipe/lib/python3.12/site-packages/accelerate/accelerator.py", line 1318, in prepare [rank1]: result = self._prepare_deepspeed(*args) [rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [rank1]: File "/conda/envs/new_pipe/lib/python3.12/site-packages/accelerate/accelerator.py", line 1815, in _prepare_deepspeed [rank1]: engine, optimizer, _, lr_scheduler = deepspeed.initialize(**kwargs) [rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [rank1]: File "/conda/envs/new_pipe/lib/python3.12/site-packages/deepspeed/__init__.py", line 193, in initialize [rank1]: engine = DeepSpeedEngine(args=args, [rank1]: ^^^^^^^^^^^^^^^^^^^^^^^^^^ [rank1]: File "/conda/envs/new_pipe/lib/python3.12/site-packages/deepspeed/runtime/engine.py", line 264, in __init__ [rank1]: self._set_distributed_vars(args) [rank1]: File "/conda/envs/new_pipe/lib/python3.12/site-packages/deepspeed/runtime/engine.py", line 1121, in _set_distributed_vars [rank1]: get_accelerator().set_device(device_rank) [rank1]: File "/conda/envs/new_pipe/lib/python3.12/site-packages/deepspeed/accelerator/cuda_accelerator.py", line 67, in set_device [rank1]: torch.cuda.set_device(device_index) [rank1]: File "/conda/envs/new_pipe/lib/python3.12/site-packages/torch/cuda/__init__.py", line 478, in set_device [rank1]: torch._C._cuda_setDevice(device) [rank1]: RuntimeError: CUDA error: invalid device ordinal [rank1]: CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. [rank1]: For debugging consider passing CUDA_LAUNCH_BLOCKING=1 [rank1]: Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
疑问
是操作参数设置有误,还是服务器性能问题?如何调整才能实现多GPU正常训练?
内容的提问来源于stack exchange,提问作者Sofia
相关产品推荐
相关产品推荐

