You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow BERT分类器出现ResourceExhaustedError的解决求助

解决Keras NLP BertClassifier训练时GPU内存不足(OOM)问题

问题场景

使用Keras NLP的BertClassifier训练模型时触发ResourceExhaustedError,提示GPU内存不足(OOM)。硬件为Nvidia RTX 3060,运行在tensorflow:latest-gpu-jupyter Docker容器中,已尝试设置TF_GPU_ALLOCATOR=cuda_malloc_async、将batch size降至1均无效。

错误日志

2024-03-22 22:53:03.932926: W external/local_tsl/tsl/framework/bfc_allocator.cc:487] Allocator (GPU_0_bfc) ran out of memory trying to allocate 192.00MiB (rounded to 201326592)requested by op StatelessRandomUniformV2
If the cause is memory fragmentation maybe the environment variable 'TF_GPU_ALLOCATOR=cuda_malloc_async' will improve the situation. 
Current allocation summary follows.
...
ResourceExhaustedError: Exception encountered when calling Dropout.call().

{{function_node __wrapped__StatelessRandomUniformV2_device_/job:localhost/replica:0/task:0/device:GPU:0}} OOM when allocating tensor with shape[128,512,768] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc [Op:StatelessRandomUniformV2] name: 

Arguments received by Dropout.call():
  • inputs=tf.Tensor(shape=(128, 512, 768), dtype=float32)
  • training=True

运行代码

import keras_nlp

classifer = keras_nlp.models.BertClassifier.from_preset(
    'bert_base_en_uncased',
    num_classes=num_classes,
    activation='softmax',
)

classifer.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'])

import os
os.environ['TF_GPU_ALLOCATOR'] = 'cuda_malloc_async'

classifer.fit(X_train, y_train, epochs=1, batch_size=128)

解决方案

  • 提前设置GPU分配器环境变量
    环境变量需要在TensorFlow/Keras初始化前设置才会生效,调整代码顺序:

    import os
    os.environ['TF_GPU_ALLOCATOR'] = 'cuda_malloc_async'
    
    import keras_nlp
    # 后续模型创建、编译、训练代码不变
    
  • 启用GPU内存增长模式
    让TensorFlow根据需求动态分配GPU内存,避免一次性占满显存:

    import tensorflow as tf
    gpus = tf.config.list_physical_devices('GPU')
    if gpus:
        try:
            for gpu in gpus:
                tf.config.experimental.set_memory_growth(gpu, True)
        except RuntimeError as e:
            print(e)
    
    # 之后再创建模型
    
  • 使用混合精度训练
    自动将部分张量转为FP16格式,大幅减少内存占用,同时不显著降低模型精度:

    from tensorflow.keras import mixed_precision
    mixed_precision.set_global_policy('mixed_float16')
    
    # 模型创建、编译代码不变
    
  • 缩短输入序列长度
    Bert默认序列长度为512,如果任务不需要这么长的序列,可在创建模型时指定更短的长度(如256或128):

    classifer = keras_nlp.models.BertClassifier.from_preset(
        'bert_base_en_uncased',
        num_classes=num_classes,
        activation='softmax',
        sequence_length=256  # 调整为适配任务的长度
    )
    
  • 改用更小的模型预设
    若bert_base仍占用内存过高,可尝试使用轻量版模型,比如bert_small_en_uncased:

    classifer = keras_nlp.models.BertClassifier.from_preset(
        'bert_small_en_uncased',
        num_classes=num_classes,
        activation='softmax',
    )
    
  • 启用梯度累积
    用极小batch size训练,累积多个batch的梯度后再更新参数,等效于大batch训练效果:

    accum_steps = 8  # 累积8个小batch的梯度再更新
    batch_size = 1
    
    classifer.compile(
        optimizer='adam',
        loss='sparse_categorical_crossentropy',
        metrics=['accuracy'],
        run_eagerly=False
    )
    
    for epoch in range(1):
        total_loss = 0.0
        accuracy = 0.0
        steps = 0
        for x_batch, y_batch in zip(X_train.batch(batch_size), y_train.batch(batch_size)):
            with tf.GradientTape() as tape:
                logits = classifer(x_batch, training=True)
                loss = classifer.compiled_loss(y_batch, logits)
            grads = tape.gradient(loss, classifer.trainable_variables)
            # 平均梯度
            for i in range(len(grads)):
                if grads[i] is not None:
                    grads[i] /= accum_steps
            classifer.optimizer.apply_gradients(zip(grads, classifer.trainable_variables))
            
            total_loss += loss.numpy()
            accuracy += classifer.compiled_metrics[0](y_batch, logits).numpy()
            steps += 1
            
            if steps % accum_steps == 0:
                print(f"Step {steps}, Loss: {total_loss/steps:.4f}, Accuracy: {accuracy/steps:.4f}")
        print(f"Epoch {epoch+1} finished, Loss: {total_loss/steps:.4f}, Accuracy: {accuracy/steps:.4f}")
    
  • 清理GPU内存碎片
    在训练前或间隙清理无效张量,减少内存碎片:

    import gc
    import tensorflow as tf
    
    gc.collect()
    tf.keras.backend.clear_session()
    

内容的提问来源于stack exchange,提问作者vmmgame

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 03:54:52