You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Colab训练CNN时遇ValueError:train_function返回空日志求因

训练CNN时触发ValueError: Empty logs的原因分析

问题描述

训练用于回归任务(预测fat_percentage)的CNN时,模型编译环节正常,但调用model.fit()触发ValueError: Unexpected result of train_function (Empty logs)错误。图片存储在Google Drive,通过独立CSV文件关联图片文件名与标签值。

错误回溯

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-1-8308471ad202> in <cell line: 89>()
     87 tf.data.experimental.enable_debug_mode()
     88 # Train the model
---> 89 model.fit(
     90     train_generator,
     91     steps_per_epoch=train_generator.samples // batch_size,

1 frames
/usr/local/lib/python3.10/dist-packages/keras/engine/training.py in fit(self, x, y, batch_size, epochs, verbose, callbacks, validation_split, validation_data, shuffle, class_weight, sample_weight, initial_epoch, steps_per_epoch, validation_steps, validation_batch_size, validation_freq, max_queue_size, workers, use_multiprocessing)
   1695                 logs = tf_utils.sync_to_numpy_or_python_type(logs)
   1696                 if logs is None:
-> 1697                     raise ValueError(
   1698                         "Unexpected result of `train_function` "
   1699                         "(Empty logs). Please use "

ValueError: Unexpected result of `train_function` (Empty logs). Please use `Model.compile(..., run_eagerly=True)`, or `tf.config.run_functions_eagerly(True)` for more information of where went wrong, or file a issue/bug to `tf.keras`.

完整代码

import tensorflow as tf
from tensorflow.keras.preprocessing.image import ImageDataGenerator
import pandas as pd
import os

from google.colab import drive
drive.mount('/content/gdrive')
%cd /content/gdrive/My Drive/

# Set the paths to your dataset directories
train_dir = './train'
test_dir = './test'

test_items = os.listdir(test_dir)
print("Items in test_dir:")
for item in test_items:
    print(item)

# Load the data into a dataframe
train_df = pd.read_csv('./train.csv')
test_df = pd.read_csv('./test.csv')

# Define image parameters
img_width, img_height = 150, 150
input_shape = (img_width, img_height, 3)  # Adjust for grayscale images

# Define hyperparameters
batch_size = 32
epochs = 10
learning_rate = 0.001

# Create data generators with data augmentation for training and validation
train_datagen = ImageDataGenerator(
    rescale=1./255,
    rotation_range=10,
    width_shift_range=0.1,
    height_shift_range=0.1,
    shear_range=0.1,
    zoom_range=0.1,
    horizontal_flip=True,
    fill_mode='nearest'
)

test_datagen = ImageDataGenerator(rescale=1./255)

train_generator = train_datagen.flow_from_dataframe(
    train_df,
    x_col='filename',
    y_col='fat_percentage',
    directory=train_dir,
    target_size=(img_width, img_height),
    batch_size=batch_size,
    class_mode='raw'
)

test_generator = test_datagen.flow_from_dataframe(
    test_df,
    x_col='filename',
    y_col='fat_percentage',
    directory=test_dir,
    target_size=(img_width, img_height),
    batch_size=batch_size,
    class_mode='raw'
)

# Define the CNN architecture
model = tf.keras.models.Sequential([
    tf.keras.layers.Conv2D(32, (3, 3), activation='relu', input_shape=input_shape),
    tf.keras.layers.MaxPooling2D((2, 2)),
    tf.keras.layers.Conv2D(64, (3, 3), activation='relu'),
    tf.keras.layers.MaxPooling2D((2, 2)),
    tf.keras.layers.Conv2D(128, (3, 3), activation='relu'),
    tf.keras.layers.MaxPooling2D((2, 2)),
    tf.keras.layers.Flatten(),
    tf.keras.layers.Dense(64, activation='relu'),
    tf.keras.layers.Dense(1)  # No activation function for regression task
])

# Compile the model
model.compile(
    loss='mse',
    optimizer=tf.keras.optimizers.Adam(learning_rate=learning_rate),
    metrics=['mae'],
)

tf.config.run_functions_eagerly(True)
tf.data.experimental.enable_debug_mode()
# Train the model
model.fit(
    train_generator,
    steps_per_epoch=train_generator.samples // batch_size,
    epochs=epochs
)

# Evaluate the model on the test set
test_loss, test_mae = model.evaluate(test_generator, verbose=2)
print(f'Test Loss: {test_loss:.4f}')
print(f'Test MAE: {test_mae:.4f}')

错误原因及排查方向

1. 数据生成器路径匹配错误

  • 核心问题:CSV中filename列的内容与train_dir下的实际文件名不匹配。比如CSV里写了train/xxx.jpg,但train_dir已经是./train,生成器会去./train/train/xxx.jpg找文件,导致找不到图片,批次数据为空。
  • 排查方法:打印train_df['filename'].head()和os.listdir(train_dir),确认文件名完全一致(Linux环境区分大小写)。

2. 标签列数据类型异常

  • 核心问题:fat_percentage列如果是字符串类型(dtype为object),flow_from_dataframe无法解析为数值型标签,导致标签数据为空。
  • 排查方法:运行print(train_df['fat_percentage'].dtype),若为object则执行train_df['fat_percentage'] = train_df['fat_percentage'].astype(float)转换类型。

3. Google Drive路径访问问题

  • 核心问题:路径中的空格(My Drive)可能导致解析异常,或者部分文件没有访问权限。
  • 排查方法:替换路径为Colab支持的无空格别名/content/gdrive/MyDrive/,并用os.listdir(train_dir)确认能正常列出所有图片。

4. steps_per_epoch计算为0

  • 核心问题:如果训练样本数小于batch_size,train_generator.samples // batch_size结果为0,导致训练没有可执行的步骤,触发空日志错误。
  • 排查方法:打印train_generator.samples和batch_size,若样本数不足,可减小batch_size,或者直接去掉steps_per_epoch参数,让Keras自动计算。

5. Eager执行模式冲突

  • 核心问题:同时启用tf.config.run_functions_eagerly(True)和tf.data.experimental.enable_debug_mode()可能导致内部逻辑冲突。
  • 解决方法:只保留其中一种调试模式,推荐直接在编译时设置:
    model.compile(
        loss='mse',
        optimizer=tf.keras.optimizers.Adam(learning_rate=learning_rate),
        metrics=['mae'],
        run_eagerly=True
    )
    

内容的提问来源于stack exchange,提问作者Axen_Rangs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 05:44:58