TensorFlow:GPU训练模型时为何出现CPU资源耗尽错误?
问题描述
- 使用GPU:NVIDIA GeForce GTX1050
- 症状:GPU训练时CPU内存持续占用直至耗尽,程序崩溃
- 报错信息:
Resource exhausted: OOM when allocating tensor with shape[64,80,80,2] and type float on /job:localhost/replica:0/task:0/device:CPU:0 by allocator
- 相关CUDA配置代码:
# CUDA # os.environ["CUDA_VISIBLE_DEVICES"] = "0, 1" assert tf.test.is_gpu_available() assert tf.test.is_built_with_cuda()
- 排查情况:任务管理器显示内存耗尽,已排除大量列表/数组/数据占用内存,且通过性能分析工具监控
- 环境版本:tensorflow-gpu:2.10.1,cudatoolkit:11.2.2
- 前置操作:重置电脑并重新安装所有环境后出现此问题
排查与解决方案
1. 修正GPU配置参数
你的GTX1050是单GPU设备,配置中CUDA_VISIBLE_DEVICES = "0, 1"会让TensorFlow尝试调用不存在的GPU,导致部分张量被迫分配到CPU。修改为:
os.environ["CUDA_VISIBLE_DEVICES"] = "0"
同时替换已弃用的tf.test.is_gpu_available(),用标准方式确认GPU识别情况:
gpus = tf.config.list_physical_devices('GPU') if gpus: print(f"识别到GPU设备: {gpus}") else: print("未识别到GPU设备")
2. 强制张量分配到GPU
显式指定模型和数据的设备上下文,避免TensorFlow自动回退到CPU:
with tf.device('/GPU:0'): # 模型定义、数据加载、训练逻辑全部放在此上下文内 model = build_your_model() model.compile(optimizer='adam', loss='sparse_categorical_crossentropy') model.fit(x_train, y_train, epochs=10)
如果使用numpy数组作为训练数据,手动转移到GPU:
x_train = tf.convert_to_tensor(x_train, dtype=tf.float32) x_train = x_train.gpu()
3. 启用GPU内存按需增长
GTX1050显存容量有限(通常2G/4G),TensorFlow默认会占用全部显存,可能导致部分操作被迫转移到CPU。开启内存增长模式:
gpus = tf.config.list_physical_devices('GPU') if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) except RuntimeError as e: print(e)
4. 校验环境依赖兼容性
TF2.10.1官方要求搭配CUDA 11.2和cuDNN 8.1.0,仅安装cudatoolkit不够:
- 卸载现有CUDA工具包,重新安装CUDA 11.2.2
- 下载对应版本的cuDNN 8.1.0,解压后将
bin、include、lib目录下的文件复制到CUDA安装目录对应位置 - 重新运行GPU识别代码,确认设备被正常识别
5. 排查训练流程中的内存泄漏
即使没有显式的大数据占用,训练过程中的中间张量、回调函数累积数据也可能导致CPU内存泄漏:
- 使用
tf.debugging.experimental.enable_dump_debug_info("./debug_log", tensor_debug_mode="FULL_HEALTH")记录张量分配位置,定位CPU上的大张量来源 - 检查
ModelCheckpoint、TensorBoard等回调函数,避免在每个epoch后保存未清理的大量临时对象
内容的提问来源于stack exchange,提问作者E.T.Tuna
相关产品推荐
相关产品推荐

