You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow中ResourceExhaustedError问题的解决求助

解决Hugging Face问答任务训练时的GPU显存不足问题

我在复现Hugging Face的《问答任务》和《问答NLP课程》内容时,执行model.fit()触发了ResourceExhaustedError(GPU显存不足OOM)。已尝试将batch_size降至16、限制GPU内存增长,但问题仍未解决。报错日志如下:

---------------------------------------------------------------------------
ResourceExhaustedError                    Traceback (most recent call last)
Cell In[14], line 1
----> 1 model.fit(x=tf_train_set, batch_size=16, validation_data=tf_validation_set, epochs=3, callbacks=[callback])
ResourceExhaustedError: Graph execution error:

Detected at node 'tf_distil_bert_for_question_answering/distilbert/transformer/layer_._4/attention/dropout_14/dropout/random_uniform/RandomUniform' defined at (most recent call last):

此处省略大量文件列表

Node: 'tf_distil_bert_for_question_answering/distilbert/transformer/layer_._4/attention/dropout_14/dropout/random_uniform/RandomUniform'
OOM when allocating tensor with shape[16,12,384,384] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc
     [[{{node tf_distil_bert_for_question_answering/distilbert/transformer/layer_._4/attention/dropout_14/dropout/random_uniform/RandomUniform}}]]
Hint: If you want to see a list of allocated tensors when OOM happens, add report_tensor_allocations_upon_oom to RunOptions for current allocation info. This isn't available when running in Eager mode.
 [Op:__inference_train_function_9297]

实用显存优化方案

  • 进一步降低batch size:尝试将batch_size设为8或4,报错中的张量形状[16,12,384,384]显示当前batch下注意力层显存占用过高,缩小batch size能直接削减显存开销
  • 启用梯度累积:用小batch size训练,累积多步梯度后再更新参数,效果等价于大batch size。示例代码:
accumulation_steps = 4  # 累积4步,等价于batch_size=64
optimizer = tf.keras.optimizers.Adam()

for epoch in range(3):
    model.train()
    total_loss = 0.0
    for step, batch in enumerate(tf_train_set):
        with tf.GradientTape() as tape:
            outputs = model(**batch)
            loss = outputs.loss
            loss = loss / accumulation_steps  # 均分损失到每一步
        
        grads = tape.gradient(loss, model.trainable_variables)
        optimizer.apply_gradients(zip(grads, model.trainable_variables))
        
        total_loss += loss.numpy() * accumulation_steps
        if (step + 1) % accumulation_steps == 0:
            optimizer.zero_grad()
    
    print(f"Epoch {epoch+1} 训练损失: {total_loss / len(tf_train_set)}")
    # 验证环节
    model.evaluate(tf_validation_set)
  • 开启混合精度训练:利用TensorFlow的混合精度减少显存占用,只需在训练前添加:
from tensorflow.keras.mixed_precision import set_global_policy
set_global_policy('mixed_float16')
  • 缩短输入序列长度:如果任务允许,将数据预处理时的max_length从默认的512调整为256,减少每个样本的显存占用
  • 手动清理显存:训练前执行以下代码释放GPU残留内存:
import tensorflow as tf
tf.keras.backend.clear_session()
gpu_devices = tf.config.list_physical_devices('GPU')
if gpu_devices:
    tf.config.experimental.set_memory_growth(gpu_devices[0], True)

内容的提问来源于stack exchange,提问作者Rumblerock

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 00:32:51