运行TensorFlow CNN触发RESOURCE_EXHAUSTED OOM错误如何解决
问题说明
运行卷积神经网络(CNN)任务时触发如下GPU显存溢出(OOM)报错:
tensorflow.python.framework.errors_impl.ResourceExhaustedError: 2 root error(s) found. (0) RESOURCE_EXHAUSTED: OOM when allocating tensor with shape[60000,32,393,2] and type float on /job:localhost/replica:0/task:0/device:GPU:1 by allocator GPU_1_bfc [[{{node sequential/conv2d_2/Relu}}]] Hint: If you want to see a list of allocated tensors when OOM happens, add report_tensor_allocations_upon_oom to RunOptions for current allocation info. This isn't available when running in Eager mode. [[StatefulPartitionedCall/_31]] Hint: If you want to see a list of allocated tensors when OOM happens, add report_tensor_allocations_upon_oom to RunOptions for current allocation info. This isn't available in Eager mode. (1) RESOURCE_EXHAUSTED: OOM when allocating tensor with shape[60000,32,393,2] and type float on /job:localhost/replica:0/task:0/device:GPU:1 by allocator GPU_1_bfc [[{{node sequential/conv2d_2/Relu}}]] Hint: If you want to see a list of allocated tensors when OOM happens, add report_tensor_allocations_upon_oom to RunOptions for current allocation info. This isn't available in Eager mode. 0 successful operations. 0 derived errors ignored. srun: error: gpu01: task 0: Exited with exit code 1
报错根因
报错出现在sequential/conv2d_2/Relu计算节点,尝试分配形状为[60000,32,393,2]的float32型张量时GPU显存不足。从张量形状可以判断,当前配置将全量60000条样本作为单个批次输入网络,单批次张量尺寸叠加模型参数、反向传播缓存的显存占用,超出了GPU可用显存上限。
解决方案
按生效优先级从高到低尝试以下操作:
- 调小batch size:这是解决该问题最直接的方案。当前配置相当于把整个数据集一次性塞进显存,把batch size下调到32、64、128这类常规取值,显存占用会随batch size等比例下降,调整后基本能解决绝大多数这类OOM问题。
- 配置显存动态增长:在代码最开头、初始化网络前加入TensorFlow显存按需分配配置,避免显存提前被占满导致分配失败,代码如下:
import tensorflow as tf gpu_devices = tf.config.list_physical_devices('GPU') for gpu in gpu_devices: tf.config.experimental.set_memory_growth(gpu, True) - 开启混合精度训练:将默认的float32计算精度切换为混合精度,大部分张量会用float16格式存储,可直接降低近一半的显存占用,开启代码如下:
tf.keras.mixed_precision.set_global_policy('mixed_float16') - 压缩张量尺寸:如果输入图像分辨率过高,可以在数据预处理阶段适当缩小输入尺寸;也可以在网络靠前位置加入池化层、适当增大卷积步长,快速降低中间特征图的高宽维度,减少后续层的显存占用;还可以适当调低卷积层的输出通道数,比如把报错层的32通道降到16,线性降低该层输出的显存占用。
- 清理冗余显存占用:训练前通过
nvidia-smi命令查看GPU上运行的进程,关闭其他不需要的、占用显存的程序;如果是在Jupyter类交互环境中运行,重启内核清空残留的无用张量。
注:报错提示中提到的
report_tensor_allocations_upon_oom参数在TensorFlow默认开启的Eager模式下不生效,无需额外添加该参数做排查。
内容的提问来源于stack exchange,提问作者rif3aa dev
相关产品推荐
相关产品推荐

