You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

VGG16迁移学习model.fit()时TensorFlow GPU内存错误求助

解决Ubuntu 24+RTX4070 Super上TensorFlow GPU内存溢出问题

问题背景

在Ubuntu 24系统的RTX 4070 Super(12GB显存)上运行VGG16迁移学习模型,调用model.fit()时触发GPU内存错误。单张224x224 RGB图像块为float32格式约602KB,理论最大batch size约5000,但实际调至1仍报错。

错误日志

2024-09-23 05:49:25.162827: I external/local_tsl/tsl/framework/bfc_allocator.cc:1112] Sum Total of in-use chunks: 3.76GiB
2024-09-23 05:49:25.162840: I external/local_tsl/tsl/framework/bfc_allocator.cc:1114] Total bytes in pool: 10671489024 memory_limit_: 10671489024 available bytes: 0 curr_region_allocation_bytes_: 21342978048
2024-09-23 05:49:25.162860: I external/local_tsl/tsl/framework/bfc_allocator.cc:1119] Stats: 
Limit:                     10671489024
InUse:                      4039777024
MaxInUse:                   4091143936
NumAllocs:                         215
MaxAllocSize:               3507001344
Reserved:                            0
PeakReserved:                        0
LargestFreeBlock:                    0

2024-09-23 05:49:25.162904: W external/local_tsl/tsl/framework/bfc_allocator.cc:499] ***************************************_____________________________________________________________
Traceback (most recent call last):
  File "/home/aiworker9/code/py/aimodels/common/prepare_data_vgg16.py", line 363, in <module>
    scores, history = model_fit_image_label_array(
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/aiworker9/code/py/aimodels/common/prepare_data_vgg16.py", line 260, in model_fit_image_label_array
    history = model.fit(
              ^^^^^^^^^^
  File "/home/aiworker9/code/py/myenv/lib/python3.12/site-packages/keras/src/utils/traceback_utils.py", line 122, in error_handler
    raise e.with_traceback(filtered_tb) from None
  File "/home/aiworker9/code/py/myenv/lib/python3.12/site-packages/tensorflow/python/framework/constant_op.py", line 108, in convert_to_eager_tensor
    return ops.EagerTensor(value, ctx.device_name, dtype)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
tensorflow.python.framework.errors_impl.InternalError: Failed copying input tensor from /job:localhost/replica:0/task:0/device:CPU:0 to /job:localhost/replica:0/task:0/device:GPU:0 in order to run _EagerConst: Dst tensor is not initialized.

已尝试的解决方法

  • 将model.fit()和data.batch的batch size调低至1;
  • 添加GPU内存管理代码:
import os
os.environ['TF_GPU_ALLOCATOR'] = 'cuda_malloc_async'  #reduce memory fragmentation.
#clear gpu memory
tf.keras.backend.clear_session()

#reduce memory footprint. nvidia gpu related
os.environ['XLA_FLAGS'] = '--xla_gpu_strict_conv_algorithm_picker=false'
from tensorflow.keras import mixed_precision
mixed_precision.set_global_policy('mixed_float16')
#reduce GPU memory
gpus = tf.config.list_physical_devices('GPU')
if gpus:
    try:
        for gpu in gpus:
            tf.config.experimental.set_memory_growth(gpu, True)
    except RuntimeError as e:
        print(e)
  • 运行nvidia-smi查看显存(错误发生前输出):
    nvidia-smi输出

模型结构(部分)

│ (Embedding)         │                  │           │                   │
├─────────────────────┼──────────────────┼───────────┼───────────────────┤
│ dropout (Dropout)   │ (None, 256)      │         0 │ dense[0][0]       │
├─────────────────────┼──────────────────┼───────────┼───────────────────┤
│ flatten_1 (Flatten) │ (None, 150)      │         0 │ embedding[0][0]   │
├─────────────────────┼──────────────────┼───────────┼───────────────────┤
│ concatenate         │ (None, 406)      │         0 │ dropout[0][0],    │
│ (Concatenate)       │                  │           │ flatten_1[0][0]   │
├─────────────────────┼──────────────────┼───────────┼───────────────────┤
│ user_output (Dense) │ (None, 3)        │     1,221 │ concatenate[0][0] │
├─────────────────────┼──────────────────┼───────────┼───────────────────┤
│ binary_output       │ (None, 2)        │       814 │ concatenate[0][0] │
│ (Dense)             │                  │           │                   │
└─────────────────────┴──────────────────┴───────────┴───────────────────┘
 Total params: 21,139,657 (80.64 MB)
 Trainable params: 6,424,969 (24.51 MB)
 Non-trainable params: 14,714,688 (56.13 MB)

数据集创建代码

def return_dataset_image_embedding(data, window_size=(224, 224), step_size=112, shuffle_buffer_size=1000, prefetch_buffer_size=1):
    # Example DataFrame
    
    print("\nBatch size: ", BATCH_SIZE)
    print("\nShuffle Buffer size : ", shuffle_buffer_size)
    print("\nPrefetch Buffer size : ", prefetch_buffer_size)
    df = pd.DataFrame(data)
    
    image_paths = './data/'+df['image_id'].values
    label_user_ids = df['label_user_id'].values
    label_binary_flags = df['label_binary_flag'].values

    image_patches, patch_label_user_ids, patch_label_binary_flags = preprocess_image_patches_image_embedding(image_paths, label_user_ids, label_binary_flags, window_size, step_size)
        
    # Create TensorFlow dataset
    patch_dataset = tf.data.Dataset.from_tensor_slices((image_patches, patch_label_user_ids, patch_label_binary_flags))
    
    # Shuffle, augment, batch, and prefetch the dataset
    patch_dataset = patch_dataset.shuffle(buffer_size=shuffle_buffer_size)  # Shuffle data
    patch_dataset = patch_dataset.map(augment_data_image_embedding)  # Apply augmentation
    patch_dataset = patch_dataset.batch(batch_size=BATCH_SIZE)  # Create batches
    patch_dataset = patch_dataset.prefetch(buffer_size=prefetch_buffer_size)  # reduce gpu memory usage
    
    return patch_dataset

解决方案

1. 修复数据集全量加载问题

你的代码中preprocess_image_patches_image_embedding会一次性生成所有图像块并加载到CPU内存,后续转GPU时极易触发内存溢出。改为流式加载处理:

def generate_patches(img, window_size=(224,224), step_size=112):
    # 实现你的图像块分割逻辑,返回shape为(num_patches, 224,224,3)的张量
    patches = []
    h, w = img.shape[0], img.shape[1]
    for y in range(0, h - window_size[0] + 1, step_size):
        for x in range(0, w - window_size[1] + 1, step_size):
            patch = img[y:y+window_size[0], x:x+window_size[1], :]
            patches.append(patch)
    return tf.stack(patches)

def load_and_process_image(image_path, label_user, label_binary, window_size=(224,224), step_size=112):
    img = tf.io.read_file(image_path)
    img = tf.image.decode_jpeg(img, channels=3)
    img = tf.image.convert_image_dtype(img, tf.float32)
    patches = generate_patches(img, window_size, step_size)
    # 为每个patch复制对应标签
    labels_user = tf.repeat(label_user, tf.shape(patches)[0])
    labels_binary = tf.repeat(label_binary, tf.shape(patches)[0])
    return patches, labels_user, labels_binary

def return_dataset_image_embedding(data, window_size=(224, 224), step_size=112, shuffle_buffer_size=1000, prefetch_buffer_size=tf.data.AUTOTUNE):
    df = pd.DataFrame(data)
    # 直接传入路径,不提前加载图像
    image_paths = tf.convert_to_tensor('./data/'+df['image_id'].values, dtype=tf.string)
    label_user_ids = tf.convert_to_tensor(df['label_user_id'].values)
    label_binary_flags = tf.convert_to_tensor(df['label_binary_flag'].values)
    
    patch_dataset = tf.data.Dataset.from_tensor_slices((image_paths, label_user_ids, label_binary_flags))
    # 按需加载处理图像,flat_map展开所有patch
    patch_dataset = patch_dataset.flat_map(lambda path, u, b: tf.data.Dataset.from_tensor_slices(load_and_process_image(path, u, b, window_size, step_size)))
    patch_dataset = patch_dataset.shuffle(shuffle_buffer_size)
    patch_dataset = patch_dataset.map(augment_data_image_embedding, num_parallel_calls=tf.data.AUTOTUNE)
    patch_dataset = patch_dataset.batch(BATCH_SIZE)
    patch_dataset = patch_dataset.prefetch(prefetch_buffer_size)
    return patch_dataset

2. 优化GPU内存配置

  • 注释掉XLA相关配置:XLA加速可能增加内存开销,尝试去掉os.environ['XLA_FLAGS'] = '--xla_gpu_strict_conv_algorithm_picker=false';
  • 强制限制GPU显存:如果内存增长失控,直接指定可用显存:
    gpus = tf.config.list_physical_devices('GPU')
    if gpus:
        tf.config.set_logical_device_configuration(
            gpus[0],
            [tf.config.LogicalDeviceConfiguration(memory_limit=10240)]  # 限制为10GB,留2GB给系统
        )
    
  • 训练前后手动清理:
    # 训练前
    tf.keras.backend.clear_session()
    
    # 训练后
    del model
    tf.keras.backend.clear_session()
    

3. 调整模型与训练细节

  • 检查数据增强函数:确保augment_data_image_embedding没有生成冗余张量,所有操作按需执行;
  • 禁用Embedding层的混合精度:Embedding层用float32更稳定,避免内存异常:
    embedding_layer = tf.keras.layers.Embedding(input_dim=..., output_dim=..., dtype='float32')
    
  • 手动指定steps_per_epoch:在model.fit()中设置该参数,避免TensorFlow自动计算时加载过多数据。

4. 系统级排查

  • 检查swap分区:确保Ubuntu swap分区至少8GB,避免CPU内存不足导致的数据传输异常;
  • 清理GPU进程:用nvidia-smi查看并关闭其他占用显存的进程,释放资源后再运行训练。

内容的提问来源于stack exchange,提问作者user938363

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 10:05:54