You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CNN训练损失停滞在51-60区间,公寓价格预测模型求助

Hey there, let's break down why your CNN model's loss is plateauing after just 1-5 epochs and walk through actionable fixes to get it converging properly.

First, Let's Diagnose the Key Issues in Your Code:

Looking at your implementation, several factors are preventing the model from learning effectively:


1. Way Too Large Convolution Kernels

Your current Conv2D layers use 32x32 and 64x64 kernels—this is way oversized for 256x256 input images. A 32x32 kernel on a 256x256 image crushes most local detail (like room layouts, fixtures, or finish quality) that's critical for price prediction. The 64x64 kernel on the subsequent layer is even worse: it would reduce the feature map to 1x1 immediately, eliminating any spatial context entirely.

Fix:

Switch to standard small kernels (3x3 or 5x5) and build up the number of filters gradually. Also, uncomment those BatchNormalization layers—they stabilize training by normalizing layer inputs:

model = Sequential([
    # First convolution block
    Conv2D(32, kernel_size=(3,3), strides=(1,1), padding='same', activation='relu', 
           input_shape=(256, 256, 3), data_format='channels_last'),
    BatchNormalization(),
    MaxPool2D(strides=2),
    # Second block
    Conv2D(64, kernel_size=(3,3), strides=(1,1), padding='same', activation='relu'),
    BatchNormalization(),
    MaxPool2D(strides=2),
    # Third block (add depth to capture more features)
    Conv2D(128, kernel_size=(3,3), strides=(1,1), padding='same', activation='relu'),
    BatchNormalization(),
    MaxPool2D(strides=2),
    # Fourth block
    Conv2D(256, kernel_size=(3,3), strides=(1,1), padding='same', activation='relu'),
    BatchNormalization(),
    MaxPool2D(strides=2),
    # Fully connected layers
    Flatten(),
    Dense(128, activation='relu', kernel_initializer='he_normal'),  # He init works better with ReLU
    Dense(64, activation='relu', kernel_initializer='he_normal'),
    Dense(1, activation='linear')
])

2. Overly High Learning Rate

Your Adam optimizer is set to lr=0.01—that's 10-100x too high for a regression task with CNNs. A high learning rate causes the model's weights to jump around the optimal solution instead of converging smoothly, leading to the loss plateau you're seeing.

Fix:

Drop the learning rate to a more reasonable range, and add a learning rate scheduler to adjust it automatically if the loss stalls:

from keras.callbacks import ReduceLROnPlateau

# Lower initial learning rate
model.compile(Adam(lr=1e-4, beta_1=0.9, beta_2=0.999), loss='mean_absolute_percentage_error')

# Add scheduler to cut learning rate in half if loss doesn't improve for 3 epochs
lr_scheduler = ReduceLROnPlateau(monitor='loss', factor=0.5, patience=3, min_lr=1e-6)

3. Test Data Loading Bug

You're loading the training dataframe for your test data—this means you're evaluating the model on the same data it trained on, which gives you meaningless results and hides real overfitting issues.

Fix:

Use test_data_df for your test generator, and avoid applying data augmentation to test data:

test_datagen = ImageDataGenerator(rescale=1./255)
test_data = test_datagen.flow_from_dataframe(
    dataframe=test_data_df, 
    x_col='filepath', 
    y_col='price', 
    class_mode='raw', 
    directory=r'C:\Users\Kojimba\PycharmProjects\DeepEval\CNN', 
    batch_size=20
)

4. Weak Data Pipeline & Training Configuration

  • No Data Augmentation: Apartment images can vary in lighting, angle, and cropping—adding augmentation helps the model generalize better and avoids early overfitting.
  • Incorrect steps_per_epoch: Hardcoding it to 24 means you might not be iterating through all your training data each epoch, which slows down learning.

Fix:

Update your data generator and training loop:

# Add data augmentation to training generator
train_datagen = ImageDataGenerator(
    rescale=1./255,
    rotation_range=10,
    width_shift_range=0.1,
    height_shift_range=0.1,
    horizontal_flip=True,
    brightness_range=[0.9, 1.1]
)

train_data = train_datagen.flow_from_dataframe(
    dataframe=train_data_df, 
    x_col='filepath', 
    y_col='price', 
    class_mode='raw', 
    directory=r'C:\Users\Kojimba\PycharmProjects\DeepEval\CNN', 
    batch_size=20
)

# Calculate steps_per_epoch properly (total training samples / batch size)
steps_per_epoch = len(train_data_df) // 20

# Train with the scheduler and validate on test data each epoch
model.fit_generator(
    train_data, 
    steps_per_epoch=steps_per_epoch, 
    epochs=100,
    validation_data=test_data,
    callbacks=[lr_scheduler]
)

5. Bonus: Tweak Loss Function & Initialization

  • Loss Function: If your apartment prices have a wide range, mean_absolute_error (MAE) might be more stable than mean_absolute_percentage_error (MAPE), which can blow up when prices are very low.
  • Initialization: Replace random_normal with he_normal for ReLU layers—it's designed to keep the variance of layer inputs consistent, preventing vanishing/exploding gradients.

Give these changes a try, and you should see the loss start to drop consistently beyond the first few epochs. Let me know if you hit any other snags!

内容的提问来源于stack exchange,提问作者Kojimba

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:53:21