VGG16迁移学习model.fit()时TensorFlow GPU内存错误求助
解决Ubuntu 24+RTX4070 Super上TensorFlow GPU内存溢出问题
问题背景
在Ubuntu 24系统的RTX 4070 Super(12GB显存)上运行VGG16迁移学习模型,调用model.fit()时触发GPU内存错误。单张224x224 RGB图像块为float32格式约602KB,理论最大batch size约5000,但实际调至1仍报错。
错误日志
2024-09-23 05:49:25.162827: I external/local_tsl/tsl/framework/bfc_allocator.cc:1112] Sum Total of in-use chunks: 3.76GiB 2024-09-23 05:49:25.162840: I external/local_tsl/tsl/framework/bfc_allocator.cc:1114] Total bytes in pool: 10671489024 memory_limit_: 10671489024 available bytes: 0 curr_region_allocation_bytes_: 21342978048 2024-09-23 05:49:25.162860: I external/local_tsl/tsl/framework/bfc_allocator.cc:1119] Stats: Limit: 10671489024 InUse: 4039777024 MaxInUse: 4091143936 NumAllocs: 215 MaxAllocSize: 3507001344 Reserved: 0 PeakReserved: 0 LargestFreeBlock: 0 2024-09-23 05:49:25.162904: W external/local_tsl/tsl/framework/bfc_allocator.cc:499] ***************************************_____________________________________________________________ Traceback (most recent call last): File "/home/aiworker9/code/py/aimodels/common/prepare_data_vgg16.py", line 363, in <module> scores, history = model_fit_image_label_array( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/aiworker9/code/py/aimodels/common/prepare_data_vgg16.py", line 260, in model_fit_image_label_array history = model.fit( ^^^^^^^^^^ File "/home/aiworker9/code/py/myenv/lib/python3.12/site-packages/keras/src/utils/traceback_utils.py", line 122, in error_handler raise e.with_traceback(filtered_tb) from None File "/home/aiworker9/code/py/myenv/lib/python3.12/site-packages/tensorflow/python/framework/constant_op.py", line 108, in convert_to_eager_tensor return ops.EagerTensor(value, ctx.device_name, dtype) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ tensorflow.python.framework.errors_impl.InternalError: Failed copying input tensor from /job:localhost/replica:0/task:0/device:CPU:0 to /job:localhost/replica:0/task:0/device:GPU:0 in order to run _EagerConst: Dst tensor is not initialized.
已尝试的解决方法
- 将
model.fit()和data.batch的batch size调低至1; - 添加GPU内存管理代码:
import os os.environ['TF_GPU_ALLOCATOR'] = 'cuda_malloc_async' #reduce memory fragmentation. #clear gpu memory tf.keras.backend.clear_session() #reduce memory footprint. nvidia gpu related os.environ['XLA_FLAGS'] = '--xla_gpu_strict_conv_algorithm_picker=false' from tensorflow.keras import mixed_precision mixed_precision.set_global_policy('mixed_float16') #reduce GPU memory gpus = tf.config.list_physical_devices('GPU') if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) except RuntimeError as e: print(e)
- 运行
nvidia-smi查看显存(错误发生前输出):
模型结构(部分)
│ (Embedding) │ │ │ │ ├─────────────────────┼──────────────────┼───────────┼───────────────────┤ │ dropout (Dropout) │ (None, 256) │ 0 │ dense[0][0] │ ├─────────────────────┼──────────────────┼───────────┼───────────────────┤ │ flatten_1 (Flatten) │ (None, 150) │ 0 │ embedding[0][0] │ ├─────────────────────┼──────────────────┼───────────┼───────────────────┤ │ concatenate │ (None, 406) │ 0 │ dropout[0][0], │ │ (Concatenate) │ │ │ flatten_1[0][0] │ ├─────────────────────┼──────────────────┼───────────┼───────────────────┤ │ user_output (Dense) │ (None, 3) │ 1,221 │ concatenate[0][0] │ ├─────────────────────┼──────────────────┼───────────┼───────────────────┤ │ binary_output │ (None, 2) │ 814 │ concatenate[0][0] │ │ (Dense) │ │ │ │ └─────────────────────┴──────────────────┴───────────┴───────────────────┘ Total params: 21,139,657 (80.64 MB) Trainable params: 6,424,969 (24.51 MB) Non-trainable params: 14,714,688 (56.13 MB)
数据集创建代码
def return_dataset_image_embedding(data, window_size=(224, 224), step_size=112, shuffle_buffer_size=1000, prefetch_buffer_size=1): # Example DataFrame print("\nBatch size: ", BATCH_SIZE) print("\nShuffle Buffer size : ", shuffle_buffer_size) print("\nPrefetch Buffer size : ", prefetch_buffer_size) df = pd.DataFrame(data) image_paths = './data/'+df['image_id'].values label_user_ids = df['label_user_id'].values label_binary_flags = df['label_binary_flag'].values image_patches, patch_label_user_ids, patch_label_binary_flags = preprocess_image_patches_image_embedding(image_paths, label_user_ids, label_binary_flags, window_size, step_size) # Create TensorFlow dataset patch_dataset = tf.data.Dataset.from_tensor_slices((image_patches, patch_label_user_ids, patch_label_binary_flags)) # Shuffle, augment, batch, and prefetch the dataset patch_dataset = patch_dataset.shuffle(buffer_size=shuffle_buffer_size) # Shuffle data patch_dataset = patch_dataset.map(augment_data_image_embedding) # Apply augmentation patch_dataset = patch_dataset.batch(batch_size=BATCH_SIZE) # Create batches patch_dataset = patch_dataset.prefetch(buffer_size=prefetch_buffer_size) # reduce gpu memory usage return patch_dataset
解决方案
1. 修复数据集全量加载问题
你的代码中preprocess_image_patches_image_embedding会一次性生成所有图像块并加载到CPU内存,后续转GPU时极易触发内存溢出。改为流式加载处理:
def generate_patches(img, window_size=(224,224), step_size=112): # 实现你的图像块分割逻辑,返回shape为(num_patches, 224,224,3)的张量 patches = [] h, w = img.shape[0], img.shape[1] for y in range(0, h - window_size[0] + 1, step_size): for x in range(0, w - window_size[1] + 1, step_size): patch = img[y:y+window_size[0], x:x+window_size[1], :] patches.append(patch) return tf.stack(patches) def load_and_process_image(image_path, label_user, label_binary, window_size=(224,224), step_size=112): img = tf.io.read_file(image_path) img = tf.image.decode_jpeg(img, channels=3) img = tf.image.convert_image_dtype(img, tf.float32) patches = generate_patches(img, window_size, step_size) # 为每个patch复制对应标签 labels_user = tf.repeat(label_user, tf.shape(patches)[0]) labels_binary = tf.repeat(label_binary, tf.shape(patches)[0]) return patches, labels_user, labels_binary def return_dataset_image_embedding(data, window_size=(224, 224), step_size=112, shuffle_buffer_size=1000, prefetch_buffer_size=tf.data.AUTOTUNE): df = pd.DataFrame(data) # 直接传入路径,不提前加载图像 image_paths = tf.convert_to_tensor('./data/'+df['image_id'].values, dtype=tf.string) label_user_ids = tf.convert_to_tensor(df['label_user_id'].values) label_binary_flags = tf.convert_to_tensor(df['label_binary_flag'].values) patch_dataset = tf.data.Dataset.from_tensor_slices((image_paths, label_user_ids, label_binary_flags)) # 按需加载处理图像,flat_map展开所有patch patch_dataset = patch_dataset.flat_map(lambda path, u, b: tf.data.Dataset.from_tensor_slices(load_and_process_image(path, u, b, window_size, step_size))) patch_dataset = patch_dataset.shuffle(shuffle_buffer_size) patch_dataset = patch_dataset.map(augment_data_image_embedding, num_parallel_calls=tf.data.AUTOTUNE) patch_dataset = patch_dataset.batch(BATCH_SIZE) patch_dataset = patch_dataset.prefetch(prefetch_buffer_size) return patch_dataset
2. 优化GPU内存配置
- 注释掉XLA相关配置:XLA加速可能增加内存开销,尝试去掉
os.environ['XLA_FLAGS'] = '--xla_gpu_strict_conv_algorithm_picker=false'; - 强制限制GPU显存:如果内存增长失控,直接指定可用显存:
gpus = tf.config.list_physical_devices('GPU') if gpus: tf.config.set_logical_device_configuration( gpus[0], [tf.config.LogicalDeviceConfiguration(memory_limit=10240)] # 限制为10GB,留2GB给系统 ) - 训练前后手动清理:
# 训练前 tf.keras.backend.clear_session() # 训练后 del model tf.keras.backend.clear_session()
3. 调整模型与训练细节
- 检查数据增强函数:确保
augment_data_image_embedding没有生成冗余张量,所有操作按需执行; - 禁用Embedding层的混合精度:Embedding层用float32更稳定,避免内存异常:
embedding_layer = tf.keras.layers.Embedding(input_dim=..., output_dim=..., dtype='float32') - 手动指定
steps_per_epoch:在model.fit()中设置该参数,避免TensorFlow自动计算时加载过多数据。
4. 系统级排查
- 检查swap分区:确保Ubuntu swap分区至少8GB,避免CPU内存不足导致的数据传输异常;
- 清理GPU进程:用
nvidia-smi查看并关闭其他占用显存的进程,释放资源后再运行训练。
内容的提问来源于stack exchange,提问作者user938363
相关产品推荐
相关产品推荐

