You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow:GPU训练模型时为何出现CPU资源耗尽错误?

问题描述
  • 使用GPU:NVIDIA GeForce GTX1050
  • 症状:GPU训练时CPU内存持续占用直至耗尽,程序崩溃
  • 报错信息:

Resource exhausted: OOM when allocating tensor with shape[64,80,80,2] and type float on /job:localhost/replica:0/task:0/device:CPU:0 by allocator

  • 相关CUDA配置代码:
# CUDA #
os.environ["CUDA_VISIBLE_DEVICES"] = "0, 1"
assert tf.test.is_gpu_available()
assert tf.test.is_built_with_cuda()
  • 排查情况:任务管理器显示内存耗尽,已排除大量列表/数组/数据占用内存,且通过性能分析工具监控
  • 环境版本:tensorflow-gpu:2.10.1,cudatoolkit:11.2.2
  • 前置操作:重置电脑并重新安装所有环境后出现此问题
排查与解决方案

1. 修正GPU配置参数

你的GTX1050是单GPU设备,配置中CUDA_VISIBLE_DEVICES = "0, 1"会让TensorFlow尝试调用不存在的GPU,导致部分张量被迫分配到CPU。修改为:

os.environ["CUDA_VISIBLE_DEVICES"] = "0"

同时替换已弃用的tf.test.is_gpu_available(),用标准方式确认GPU识别情况:

gpus = tf.config.list_physical_devices('GPU')
if gpus:
    print(f"识别到GPU设备: {gpus}")
else:
    print("未识别到GPU设备")

2. 强制张量分配到GPU

显式指定模型和数据的设备上下文,避免TensorFlow自动回退到CPU:

with tf.device('/GPU:0'):
    # 模型定义、数据加载、训练逻辑全部放在此上下文内
    model = build_your_model()
    model.compile(optimizer='adam', loss='sparse_categorical_crossentropy')
    model.fit(x_train, y_train, epochs=10)

如果使用numpy数组作为训练数据,手动转移到GPU:

x_train = tf.convert_to_tensor(x_train, dtype=tf.float32)
x_train = x_train.gpu()

3. 启用GPU内存按需增长

GTX1050显存容量有限(通常2G/4G),TensorFlow默认会占用全部显存,可能导致部分操作被迫转移到CPU。开启内存增长模式:

gpus = tf.config.list_physical_devices('GPU')
if gpus:
    try:
        for gpu in gpus:
            tf.config.experimental.set_memory_growth(gpu, True)
    except RuntimeError as e:
        print(e)

4. 校验环境依赖兼容性

TF2.10.1官方要求搭配CUDA 11.2和cuDNN 8.1.0,仅安装cudatoolkit不够:

  • 卸载现有CUDA工具包,重新安装CUDA 11.2.2
  • 下载对应版本的cuDNN 8.1.0,解压后将bin、include、lib目录下的文件复制到CUDA安装目录对应位置
  • 重新运行GPU识别代码,确认设备被正常识别

5. 排查训练流程中的内存泄漏

即使没有显式的大数据占用,训练过程中的中间张量、回调函数累积数据也可能导致CPU内存泄漏:

  • 使用tf.debugging.experimental.enable_dump_debug_info("./debug_log", tensor_debug_mode="FULL_HEALTH")记录张量分配位置,定位CPU上的大张量来源
  • 检查ModelCheckpoint、TensorBoard等回调函数,避免在每个epoch后保存未清理的大量临时对象

内容的提问来源于stack exchange,提问作者E.T.Tuna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 03:15:50