You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras分布式学习下多GPU无法分配大张量的问题排查

为何使用TensorFlow MirroredStrategy多GPU仍出现OOM错误?

问题详情

我采用TensorFlow分布式学习,执行代码如下:

os.environ["CUDA_DEVICE_ORDER"]="PCI_BUS_ID"   
os.environ["CUDA_VISIBLE_DEVICES"]="0,1,2,3"

strategy = tf.distribute.MirroredStrategy()
with strategy.scope():
    model = Basic_Model()
    model.compile(loss='mean_squared_error', optimizer=rms, metrics=['mean_squared_error'])

系统配备4张32GB的Tesla V100 GPU,nvidia-smi输出:

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 418.87.01    Driver Version: 418.87.01    CUDA Version: 10.1     |
|-------------------------------+----------------------+----------------------+ 
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|===============================+======================+======================|
|   0  Tesla V100-SXM2...  On   | 00000004:04:00.0 Off |                    0 |
| N/A   37C    P0    65W / 300W |      0MiB / 32480MiB |      0%      Default |
+-------------------------------+----------------------+----------------------+
|   1  Tesla V100-SXM2...  On   | 00000004:05:00.0 Off |                    0 |
| N/A   38C    P0    40W / 300W |      0MiB / 32480MiB |      0%      Default |
+-------------------------------+----------------------+----------------------+
|   2  Tesla V100-SXM2...  On   | 00000035:03:00.0 Off |                    0 |
| N/A   33C    P0    40W / 300W |      0MiB / 32480MiB |      0%      Default |
+-------------------------------+----------------------+----------------------+
|   3  Tesla V100-SXM2...  On   | 00000035:04:00.0 Off |                    0 |
| N/A   39C    P0    41W / 300W |      0MiB / 32480MiB |      0%      Default |
+-------------------------------+----------------------+----------------------+

+-----------------------------------------------------------------------------+
| Processes:                                                       GPU Memory |
|  GPU       PID   Type   Process name                             Usage      |
|=============================================================================|
|  No running processes found                                                 |
+-----------------------------------------------------------------------------+

运行脚本创建模型时触发错误:

tensorflow.python.framework.errors_impl.ResourceExhaustedError: OOM when allocating tensor with shape [131072,65536] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc [Op:RandomUniform]

该float类型张量需占用约34.35GB显存,系统总显存达128GB,为何无法分配?


核心原因

MirroredStrategy的核心机制是在每个GPU上完整复制一份模型实例,通过AllReduce操作同步各GPU的梯度,它不会自动将单个大张量拆分到多个GPU上。你遇到的问题本质是:这个34GB的大张量需要完整放入单张GPU的显存,但单张V100只有32GB,哪怕总显存足够,单GPU也无法容纳该张量,因此触发OOM。


解决办法

  1. 拆分大张量/模型层
    检查Basic_Model中生成该大张量的模块,手动将大张量拆分为多个小张量分片,或者将对应的模型层拆分到不同GPU上执行计算,实现模型并行。

  2. 改用适配的分布式策略
    如果需要处理单GPU装不下的大张量,可使用tf.distribute.experimental.ModelParallelism实现显式的模型并行,或者采用tf.distribute.experimental.ParameterServerStrategy将参数拆分到不同节点存储。

  3. 压缩张量体积

  • 改用低精度数据类型:将float32改为float16,可将该张量的显存占用降至约17GB,单张32GB GPU即可容纳。
  • 减少特征维度:如果业务逻辑允许,通过降维、特征选择等方式缩小张量的shape规模。
  1. 显存优化配置
    启用TensorFlow的显存增长模式,避免进程启动时一次性占满GPU显存,为张量分配预留空间:
gpus = tf.config.experimental.list_physical_devices('GPU')
for gpu in gpus:
    tf.config.experimental.set_memory_growth(gpu, True)

内容的提问来源于stack exchange,提问作者psj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 10:48:22