You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CNN训练异常求助:偶数轮次数据耗尽中断+初始准确率异常

CNN训练问题排查与解决

问题描述

  • 奇数轮次(第1、3、5轮等)训练正常,偶数轮次触发错误:Your input ran out of data; interrupting training. Make sure that your dataset or generator can generate at least steps_per_epoch * epochs batches.
  • 首个epoch准确率从1.0开始缓慢下降,而非从0逐步提升(准确率变化图:准确率变化)

已尝试操作

  • 将steps_per_epoch从len(train_generator)改为len(train_data) // batch_size,问题未解决
  • 数据集含约900张图片,排除数据量过小因素

代码片段

import os 
import tensorflow as tf
import pandas as pd
import numpy as np
from PIL import Image
import matplotlib.pyplot as plt
import urllib.request
from sklearn.model_selection import train_test_split
url = "https://raw.githubusercontent.com/mrdbourke/tensorflow-deep-learning/main/extras/helper_functions.py"
filename = "helper_functions.py"
from tensorflow.keras.preprocessing.image import ImageDataGenerator
from helper_functions import create_tensorboard_callback, plot_loss_curves, unzip_data, compare_historys, walk_through_dir, pred_and_plot 
project_path = "C:\\Users\\immaf\\Desktop\\Extended essay" 

images = []
labels = []
for subfolder in os.listdir(project_path):
    print(f"Processing folder: {subfolder}")  # Print folder name for debugging
    
    subfolder_path = os.path.join(project_path, subfolder)
    if not os.path.isdir(subfolder_path):
        continue

    # List images in the folder
    for image_filename in os.listdir(subfolder_path):
        image_path = os.path.join(subfolder_path, image_filename)
        images.append(image_path)
        labels.append(subfolder)

# Create DataFrame
data = pd.DataFrame({'image': images, 'label': labels})
pd.set_option('display.max_rows', 10)
# Split the data into training (70%), validation (15%), and test (15%)
train_data, temp_data = train_test_split(data, test_size=0.3, random_state=42, stratify=data['label'])
val_data, test_data = train_test_split(temp_data, test_size=0.5, random_state=42, stratify=temp_data['label'])
batch_size = 32
image_size = (256, 256)
# Training data generator with augmentation
train_datagen = ImageDataGenerator(
    rescale=1.0/255,  # Rescale pixel values to [0, 1]
    rotation_range=20,  # Randomly rotate images
    width_shift_range=0.2,  # Randomly shift images horizontally
    height_shift_range=0.2,  # Randomly shift images vertically
    shear_range=0.2,  # Shear angle in counter-clockwise direction
    zoom_range=0.2,  # Randomly zoom into images
    horizontal_flip=True,  # Randomly flip images
    fill_mode='nearest'  # Fill in newly created pixels
)


# Validation and test data generator (only rescaling)
val_datagen = ImageDataGenerator(rescale=1.0/255)
test_datagen = ImageDataGenerator(rescale=1.0/255)


# Create generators for training, validation, and test sets
train_generator = train_datagen.flow_from_dataframe(
    dataframe=train_data,
    x_col='image',
    y_col='label',
    target_size=image_size,
    batch_size=batch_size,
    class_mode='categorical',  # Use 'categorical' for multi-class classification
    shuffle=True,
)
val_generator = val_datagen.flow_from_dataframe(
    dataframe=val_data,
    x_col='image',
    y_col='label',
    target_size=image_size,
    batch_size=batch_size,
    class_mode='categorical',
)

test_generator = test_datagen.flow_from_dataframe(
    dataframe=test_data,
    x_col='image',
    y_col='label',
    target_size=image_size,
    batch_size=batch_size,
    class_mode='categorical'
)
from tensorflow.keras import models, layers
learning_rate = 0.001  # Adjust learning rate
filter_size = (3, 3)  # Size of the convolution filters
num_filters = 32  # Number of filters in the first Conv2D layer
dropout_rate = 0.5  # Dropout rate
pooling_layer_type = 'max'  # 'max' or 'average' pooling
strides = (1, 1)  # Stride for the convolutional layers

# Create the model
model = models.Sequential()

# First convolutional layer
model.add(layers.Conv2D(num_filters, filter_size, strides=strides, activation='relu', input_shape=(256, 256, 3)))

# Add pooling layer
if pooling_layer_type == 'max':
    model.add(layers.MaxPooling2D(pool_size=(2, 2)))
elif pooling_layer_type == 'average':
    model.add(layers.AveragePooling2D(pool_size=(2, 2)))

# Second convolutional layer
model.add(layers.Conv2D(num_filters * 2, filter_size, strides=strides, activation='relu'))  # Double the filters for deeper layers
if pooling_layer_type == 'max':
    model.add(layers.MaxPooling2D(pool_size=(2, 2)))
elif pooling_layer_type == 'average':
    model.add(layers.AveragePooling2D(pool_size=(2, 2)))

# Third convolutional layer
model.add(layers.Conv2D(num_filters * 4, filter_size, strides=strides, activation='relu'))  # Increase filters
if pooling_layer_type == 'max':
    model.add(layers.MaxPooling2D(pool_size=(2, 2)))
elif pooling_layer_type == 'average':
    model.add(layers.AveragePooling2D(pool_size=(2, 2)))

# Flattening the output
model.add(layers.Flatten())

# Fully connected layer
model.add(layers.Dense(128, activation='relu'))

# Dropout layer
model.add(layers.Dropout(dropout_rate))

# Output layer
# Change from binary classification to multi-class classification
num_classes = 2
model.add(layers.Dense(num_classes, activation='softmax'))  # Replace num_classes with the actual number of classes


# Compile the model with binary crossentropy with specified learning rate
model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=learning_rate),
              loss='categorical_crossentropy',  # Change to binary crossentropy
              metrics=['accuracy'])


# Model summary to check the structure
model.summary()
print(f"Total training samples: {len(train_data)}")
print(f"Batch size: {batch_size}")
print(f"Steps per epoch: {len(train_data) // batch_size}")


history = model.fit(
    train_generator,
    steps_per_epoch= (len(train_data) + batch_size - 1) // batch_size,
    validation_data=val_generator,
    validation_steps=len(val_generator),
    epochs=20,
    callbacks=[]
)

解决方案

针对问题1:偶数轮次数据耗尽错误

核心原因:ImageDataGenerator生成器不会自动重置数据指针,第一轮结束后指针停在最后一个batch位置,第二轮开始时无足够剩余数据;同时手动设置的步数参数可能与生成器实际可提供的batch数不匹配。

修复步骤:

  1. 给验证集生成器添加shuffle=True,确保每轮验证数据随机:
val_generator = val_datagen.flow_from_dataframe(
    dataframe=val_data,
    x_col='image',
    y_col='label',
    target_size=image_size,
    batch_size=batch_size,
    class_mode='categorical',
    shuffle=True  # 添加此行
)
  1. 移除model.fit()中的steps_per_epoch和validation_steps参数,让Keras自动根据样本量和batch_size计算步数:
history = model.fit(
    train_generator,
    validation_data=val_generator,
    epochs=20,
    callbacks=[]
)

若必须手动设置步数,需通过自定义回调在每轮前重置生成器:

class ResetGeneratorCallback(tf.keras.callbacks.Callback):
    def on_epoch_begin(self, epoch, logs=None):
        train_generator.reset()
        val_generator.reset()

# 在fit中添加回调
history = model.fit(
    train_generator,
    steps_per_epoch= (len(train_data) + batch_size - 1) // batch_size,
    validation_data=val_generator,
    validation_steps=(len(val_data) + batch_size -1) // batch_size,
    epochs=20,
    callbacks=[ResetGeneratorCallback()]
)

针对问题2:初始准确率为1.0并下降

核心原因:模型初始权重恰好使第一个batch的预测完全匹配标签,或标签映射/数据加载存在异常。

排查与修复步骤:

  1. 检查标签映射,确认类别与标签对应正确:
print(train_generator.class_indices)
  1. 验证第一个batch的预测与真实标签是否完全一致:
x_batch, y_batch = train_generator.next()
preds = model.predict(x_batch)
print("真实标签:", np.argmax(y_batch, axis=1))
print("预测结果:", np.argmax(preds, axis=1))

若完全一致,可调整输出层初始化方式:

model.add(layers.Dense(num_classes, activation='softmax', kernel_initializer='he_normal'))
  1. 检查训练集标签分布,确认包含两类样本且无重复/错误标签。

内容的提问来源于stack exchange,提问作者Adam Frank

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 04:10:54