在Colab训练CNN时遇ValueError:train_function返回空日志求因
训练CNN时触发ValueError: Empty logs的原因分析
问题描述
训练用于回归任务(预测fat_percentage)的CNN时,模型编译环节正常,但调用model.fit()触发ValueError: Unexpected result of train_function (Empty logs)错误。图片存储在Google Drive,通过独立CSV文件关联图片文件名与标签值。
错误回溯
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) <ipython-input-1-8308471ad202> in <cell line: 89>() 87 tf.data.experimental.enable_debug_mode() 88 # Train the model ---> 89 model.fit( 90 train_generator, 91 steps_per_epoch=train_generator.samples // batch_size, 1 frames /usr/local/lib/python3.10/dist-packages/keras/engine/training.py in fit(self, x, y, batch_size, epochs, verbose, callbacks, validation_split, validation_data, shuffle, class_weight, sample_weight, initial_epoch, steps_per_epoch, validation_steps, validation_batch_size, validation_freq, max_queue_size, workers, use_multiprocessing) 1695 logs = tf_utils.sync_to_numpy_or_python_type(logs) 1696 if logs is None: -> 1697 raise ValueError( 1698 "Unexpected result of `train_function` " 1699 "(Empty logs). Please use " ValueError: Unexpected result of `train_function` (Empty logs). Please use `Model.compile(..., run_eagerly=True)`, or `tf.config.run_functions_eagerly(True)` for more information of where went wrong, or file a issue/bug to `tf.keras`.
完整代码
import tensorflow as tf from tensorflow.keras.preprocessing.image import ImageDataGenerator import pandas as pd import os from google.colab import drive drive.mount('/content/gdrive') %cd /content/gdrive/My Drive/ # Set the paths to your dataset directories train_dir = './train' test_dir = './test' test_items = os.listdir(test_dir) print("Items in test_dir:") for item in test_items: print(item) # Load the data into a dataframe train_df = pd.read_csv('./train.csv') test_df = pd.read_csv('./test.csv') # Define image parameters img_width, img_height = 150, 150 input_shape = (img_width, img_height, 3) # Adjust for grayscale images # Define hyperparameters batch_size = 32 epochs = 10 learning_rate = 0.001 # Create data generators with data augmentation for training and validation train_datagen = ImageDataGenerator( rescale=1./255, rotation_range=10, width_shift_range=0.1, height_shift_range=0.1, shear_range=0.1, zoom_range=0.1, horizontal_flip=True, fill_mode='nearest' ) test_datagen = ImageDataGenerator(rescale=1./255) train_generator = train_datagen.flow_from_dataframe( train_df, x_col='filename', y_col='fat_percentage', directory=train_dir, target_size=(img_width, img_height), batch_size=batch_size, class_mode='raw' ) test_generator = test_datagen.flow_from_dataframe( test_df, x_col='filename', y_col='fat_percentage', directory=test_dir, target_size=(img_width, img_height), batch_size=batch_size, class_mode='raw' ) # Define the CNN architecture model = tf.keras.models.Sequential([ tf.keras.layers.Conv2D(32, (3, 3), activation='relu', input_shape=input_shape), tf.keras.layers.MaxPooling2D((2, 2)), tf.keras.layers.Conv2D(64, (3, 3), activation='relu'), tf.keras.layers.MaxPooling2D((2, 2)), tf.keras.layers.Conv2D(128, (3, 3), activation='relu'), tf.keras.layers.MaxPooling2D((2, 2)), tf.keras.layers.Flatten(), tf.keras.layers.Dense(64, activation='relu'), tf.keras.layers.Dense(1) # No activation function for regression task ]) # Compile the model model.compile( loss='mse', optimizer=tf.keras.optimizers.Adam(learning_rate=learning_rate), metrics=['mae'], ) tf.config.run_functions_eagerly(True) tf.data.experimental.enable_debug_mode() # Train the model model.fit( train_generator, steps_per_epoch=train_generator.samples // batch_size, epochs=epochs ) # Evaluate the model on the test set test_loss, test_mae = model.evaluate(test_generator, verbose=2) print(f'Test Loss: {test_loss:.4f}') print(f'Test MAE: {test_mae:.4f}')
错误原因及排查方向
1. 数据生成器路径匹配错误
- 核心问题:CSV中
filename列的内容与train_dir下的实际文件名不匹配。比如CSV里写了train/xxx.jpg,但train_dir已经是./train,生成器会去./train/train/xxx.jpg找文件,导致找不到图片,批次数据为空。 - 排查方法:打印
train_df['filename'].head()和os.listdir(train_dir),确认文件名完全一致(Linux环境区分大小写)。
2. 标签列数据类型异常
- 核心问题:
fat_percentage列如果是字符串类型(dtype为object),flow_from_dataframe无法解析为数值型标签,导致标签数据为空。 - 排查方法:运行
print(train_df['fat_percentage'].dtype),若为object则执行train_df['fat_percentage'] = train_df['fat_percentage'].astype(float)转换类型。
3. Google Drive路径访问问题
- 核心问题:路径中的空格(
My Drive)可能导致解析异常,或者部分文件没有访问权限。 - 排查方法:替换路径为Colab支持的无空格别名
/content/gdrive/MyDrive/,并用os.listdir(train_dir)确认能正常列出所有图片。
4. steps_per_epoch计算为0
- 核心问题:如果训练样本数小于batch_size,
train_generator.samples // batch_size结果为0,导致训练没有可执行的步骤,触发空日志错误。 - 排查方法:打印
train_generator.samples和batch_size,若样本数不足,可减小batch_size,或者直接去掉steps_per_epoch参数,让Keras自动计算。
5. Eager执行模式冲突
- 核心问题:同时启用
tf.config.run_functions_eagerly(True)和tf.data.experimental.enable_debug_mode()可能导致内部逻辑冲突。 - 解决方法:只保留其中一种调试模式,推荐直接在编译时设置:
model.compile( loss='mse', optimizer=tf.keras.optimizers.Adam(learning_rate=learning_rate), metrics=['mae'], run_eagerly=True )
内容的提问来源于stack exchange,提问作者Axen_Rangs
相关产品推荐
相关产品推荐

