You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Horovod多GPU训练CNN-LSTM模型时遇GPU ID无效报错

问题描述

使用Horovod进行多GPU分布式训练时,单GPU和双GPU运行简单卷积神经网络(CNN)均正常,但切换为CNN与LSTM结合的模型后,单GPU运行正常,双GPU启动时报错。

完整错误回溯

[1,1]<stderr>:Traceback (most recent call last):
[1,1]<stderr>:  File "horovod-PAMAP2.py", line 30, in <module>
[1,1]<stderr>:    K.set_session(tf.compat.v1.Session(config=config))
[1,1]<stderr>:  File "/usr/local/lib/python3.6/dist-packages/tensorflow/python/client/session.py", line 1586, in __init__
[1,1]<stderr>:    super(Session, self).__init__(target, graph, config=config)
[1,1]<stderr>:  File "/usr/local/lib/python3.6/dist-packages/tensorflow/python/client/session.py", line 701, in __init__
[1,1]<stderr>:    self._session = tf_session.TF_NewSessionRef(self._graph._c_graph, opts)
[1,1]<stderr>:tensorflow.python.framework.errors_impl.InvalidArgumentError: 'visible_device_list' listed an invalid GPU id '1' but visible device count is 1
[1,0]<stderr>:2022-08-14 19:21:07.506043: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x5fb58a0 initialized for platform CUDA (this does not guarantee that XLA will be used). Devices:
[1,0]<stderr>:2022-08-14 19:21:07.506129: I tensorflow/compiler/xla/service/service.cc:176]   StreamExecutor device (0): NVIDIA GeForce RTX 2080 Ti, Compute Capability 7.5
[1,0]<stderr>:2022-08-14 19:21:07.507537: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1561] Found device 0 with properties:
[1,0]<stderr>:pciBusID: 0000:18:00.0 name: NVIDIA GeForce RTX 2080 Ti computeCapability: 7.5
[1,0]<stderr>:coreClock: 1.545GHz coreCount: 68 deviceMemorySize: 10.76GiB deviceMemoryBandwidth: 573.69GiB/s
[1,0]<stderr>:2022-08-14 19:21:07.515280: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library libcudart.so.10.1
[1,0]<stderr>:2022-08-14 19:21:07.534658: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library libcublas.so.10
[1,0]<stderr>:2022-08-14 19:21:07.670187: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library libcufft.so.10
-------------------------------------------------------
Primary job  terminated normally, but 1 process returned
a non-zero exit code.. Per user-direction, the job has been aborted.
-------------------------------------------------------
--------------------------------------------------------------------------
mpirun detected that one or more processes exited with non-zero status, thus causing
the job to be terminated. The first process to do so was:

  Process name: [[7328,1],1]
  Exit code:    1
--------------------------------------------------------------------------

双GPU运行命令

/usr/local/bin/horovodrun -np 2 -H localhost:2 /usr/bin/python3 horovod-PAMAP2.py 256

内容的提问来源于stack exchange,提问作者Ahmad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 11:39:35