TensorFlow BERT分类器出现ResourceExhaustedError的解决求助
解决Keras NLP BertClassifier训练时GPU内存不足(OOM)问题
问题场景
使用Keras NLP的BertClassifier训练模型时触发ResourceExhaustedError,提示GPU内存不足(OOM)。硬件为Nvidia RTX 3060,运行在tensorflow:latest-gpu-jupyter Docker容器中,已尝试设置TF_GPU_ALLOCATOR=cuda_malloc_async、将batch size降至1均无效。
错误日志
2024-03-22 22:53:03.932926: W external/local_tsl/tsl/framework/bfc_allocator.cc:487] Allocator (GPU_0_bfc) ran out of memory trying to allocate 192.00MiB (rounded to 201326592)requested by op StatelessRandomUniformV2 If the cause is memory fragmentation maybe the environment variable 'TF_GPU_ALLOCATOR=cuda_malloc_async' will improve the situation. Current allocation summary follows. ... ResourceExhaustedError: Exception encountered when calling Dropout.call(). {{function_node __wrapped__StatelessRandomUniformV2_device_/job:localhost/replica:0/task:0/device:GPU:0}} OOM when allocating tensor with shape[128,512,768] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc [Op:StatelessRandomUniformV2] name: Arguments received by Dropout.call(): • inputs=tf.Tensor(shape=(128, 512, 768), dtype=float32) • training=True
运行代码
import keras_nlp classifer = keras_nlp.models.BertClassifier.from_preset( 'bert_base_en_uncased', num_classes=num_classes, activation='softmax', ) classifer.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy']) import os os.environ['TF_GPU_ALLOCATOR'] = 'cuda_malloc_async' classifer.fit(X_train, y_train, epochs=1, batch_size=128)
解决方案
提前设置GPU分配器环境变量
环境变量需要在TensorFlow/Keras初始化前设置才会生效,调整代码顺序:import os os.environ['TF_GPU_ALLOCATOR'] = 'cuda_malloc_async' import keras_nlp # 后续模型创建、编译、训练代码不变启用GPU内存增长模式
让TensorFlow根据需求动态分配GPU内存,避免一次性占满显存:import tensorflow as tf gpus = tf.config.list_physical_devices('GPU') if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) except RuntimeError as e: print(e) # 之后再创建模型使用混合精度训练
自动将部分张量转为FP16格式,大幅减少内存占用,同时不显著降低模型精度:from tensorflow.keras import mixed_precision mixed_precision.set_global_policy('mixed_float16') # 模型创建、编译代码不变缩短输入序列长度
Bert默认序列长度为512,如果任务不需要这么长的序列,可在创建模型时指定更短的长度(如256或128):classifer = keras_nlp.models.BertClassifier.from_preset( 'bert_base_en_uncased', num_classes=num_classes, activation='softmax', sequence_length=256 # 调整为适配任务的长度 )改用更小的模型预设
若bert_base仍占用内存过高,可尝试使用轻量版模型,比如bert_small_en_uncased:classifer = keras_nlp.models.BertClassifier.from_preset( 'bert_small_en_uncased', num_classes=num_classes, activation='softmax', )启用梯度累积
用极小batch size训练,累积多个batch的梯度后再更新参数,等效于大batch训练效果:accum_steps = 8 # 累积8个小batch的梯度再更新 batch_size = 1 classifer.compile( optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'], run_eagerly=False ) for epoch in range(1): total_loss = 0.0 accuracy = 0.0 steps = 0 for x_batch, y_batch in zip(X_train.batch(batch_size), y_train.batch(batch_size)): with tf.GradientTape() as tape: logits = classifer(x_batch, training=True) loss = classifer.compiled_loss(y_batch, logits) grads = tape.gradient(loss, classifer.trainable_variables) # 平均梯度 for i in range(len(grads)): if grads[i] is not None: grads[i] /= accum_steps classifer.optimizer.apply_gradients(zip(grads, classifer.trainable_variables)) total_loss += loss.numpy() accuracy += classifer.compiled_metrics[0](y_batch, logits).numpy() steps += 1 if steps % accum_steps == 0: print(f"Step {steps}, Loss: {total_loss/steps:.4f}, Accuracy: {accuracy/steps:.4f}") print(f"Epoch {epoch+1} finished, Loss: {total_loss/steps:.4f}, Accuracy: {accuracy/steps:.4f}")清理GPU内存碎片
在训练前或间隙清理无效张量,减少内存碎片:import gc import tensorflow as tf gc.collect() tf.keras.backend.clear_session()
内容的提问来源于stack exchange,提问作者vmmgame
相关产品推荐
相关产品推荐

