You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行TensorFlow CNN触发RESOURCE_EXHAUSTED OOM错误如何解决

问题说明

运行卷积神经网络(CNN)任务时触发如下GPU显存溢出(OOM)报错:

tensorflow.python.framework.errors_impl.ResourceExhaustedError: 2 root error(s) found.
  (0) RESOURCE_EXHAUSTED: OOM when allocating tensor with shape[60000,32,393,2] and type float on /job:localhost/replica:0/task:0/device:GPU:1 by allocator GPU_1_bfc
     [[{{node sequential/conv2d_2/Relu}}]]
Hint: If you want to see a list of allocated tensors when OOM happens, add report_tensor_allocations_upon_oom to RunOptions for current allocation info. This isn't available when running in Eager mode.

     [[StatefulPartitionedCall/_31]]
Hint: If you want to see a list of allocated tensors when OOM happens, add report_tensor_allocations_upon_oom to RunOptions for current allocation info. This isn't available in Eager mode.

  (1) RESOURCE_EXHAUSTED: OOM when allocating tensor with shape[60000,32,393,2] and type float on /job:localhost/replica:0/task:0/device:GPU:1 by allocator GPU_1_bfc
     [[{{node sequential/conv2d_2/Relu}}]]
Hint: If you want to see a list of allocated tensors when OOM happens, add report_tensor_allocations_upon_oom to RunOptions for current allocation info. This isn't available in Eager mode.

0 successful operations.
0 derived errors ignored.
srun: error: gpu01: task 0: Exited with exit code 1
报错根因

报错出现在sequential/conv2d_2/Relu计算节点,尝试分配形状为[60000,32,393,2]的float32型张量时GPU显存不足。从张量形状可以判断,当前配置将全量60000条样本作为单个批次输入网络,单批次张量尺寸叠加模型参数、反向传播缓存的显存占用,超出了GPU可用显存上限。

解决方案

按生效优先级从高到低尝试以下操作:

  • 调小batch size:这是解决该问题最直接的方案。当前配置相当于把整个数据集一次性塞进显存,把batch size下调到32、64、128这类常规取值,显存占用会随batch size等比例下降,调整后基本能解决绝大多数这类OOM问题。
  • 配置显存动态增长:在代码最开头、初始化网络前加入TensorFlow显存按需分配配置,避免显存提前被占满导致分配失败,代码如下:
    import tensorflow as tf
    gpu_devices = tf.config.list_physical_devices('GPU')
    for gpu in gpu_devices:
        tf.config.experimental.set_memory_growth(gpu, True)
    
  • 开启混合精度训练:将默认的float32计算精度切换为混合精度,大部分张量会用float16格式存储,可直接降低近一半的显存占用,开启代码如下:
    tf.keras.mixed_precision.set_global_policy('mixed_float16')
    
  • 压缩张量尺寸:如果输入图像分辨率过高,可以在数据预处理阶段适当缩小输入尺寸;也可以在网络靠前位置加入池化层、适当增大卷积步长,快速降低中间特征图的高宽维度,减少后续层的显存占用;还可以适当调低卷积层的输出通道数,比如把报错层的32通道降到16,线性降低该层输出的显存占用。
  • 清理冗余显存占用:训练前通过nvidia-smi命令查看GPU上运行的进程,关闭其他不需要的、占用显存的程序;如果是在Jupyter类交互环境中运行,重启内核清空残留的无用张量。

注:报错提示中提到的report_tensor_allocations_upon_oom参数在TensorFlow默认开启的Eager模式下不生效,无需额外添加该参数做排查。

内容的提问来源于stack exchange,提问作者rif3aa dev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 22:27:32