You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CNN合金显微图像抗拉强度回归预测值趋同及RMSE优化咨询

合金显微组织图像抗拉强度预测问题

实验背景

本次实验输入为合金显微组织图像,预测目标为合金的抗拉强度(TS)。将图像输入搭建的CNN回归模型后,未得到预期的分散预测值,实际输出预测结果几乎全部为相近数值,且RMSE表现极差。

实验代码

import numpy as np
import pandas as pd
from pathlib import Path
import os.path

from sklearn.model_selection import train_test_split

import tensorflow as tf

from sklearn.metrics import r2_score


from keras.applications.efficientnet import EfficientNetB3
import gc
from keras.models import Sequential
from keras import layers, models

from keras import Input
from keras.models import Model
from keras.preprocessing.image import ImageDataGenerator
from keras import optimizers, initializers, regularizers, metrics
from keras.callbacks import ModelCheckpoint
import os
from glob import glob
from PIL import Image
import matplotlib.pyplot as plt
import numpy as np
from tensorflow.keras import optimizers
from keras.layers import Conv2D,MaxPool2D,GlobalAveragePooling2D,AveragePooling2D
from keras.layers import Dense,Dropout,Activation,Flatten
import sys
# Repository source: https://github.com/qubvel/efficientnet
sys.path.append(os.path.abspath('../input/efficientnet/efficientnet-master/efficientnet-master/'))
from tensorflow.keras.callbacks import EarlyStopping, ReduceLROnPlateau


image_dir = Path('/content/drive/MyDrive/processed')

filepaths = pd.Series(list(image_dir.glob(r'**/*.jpg')), name='Filepath').astype(str) 

TS = pd.Series(sorted([int(l.split('TS_')[1].split('/pre')[0]) for l in filepaths]),name='TS').astype(np.int)
images = pd.concat([filepaths, TS], axis=1).sample(frac=1.0, 
random_state=1).reset_index(drop=True)


image_df = images.sample(2020, random_state=1).reset_index(drop=True)

train_df, test_df = train_test_split(image_df, train_size=0.7, shuffle=True, random_state=1)

train_generator = tf.keras.preprocessing.image.ImageDataGenerator(
    rescale=1./255,
    validation_split=0.2
)

test_generator = tf.keras.preprocessing.image.ImageDataGenerator(
    rescale=1./255
)

train_input = train_generator.flow_from_dataframe(
    dataframe=train_df,
    x_col='Filepath',
    y_col='TS',
    target_size=(256, 256),
    color_mode='grayscale',
    class_mode='raw',
    batch_size=1,
    shuffle=True,
    seed=42,
    subset='training'
)

val_input = train_generator.flow_from_dataframe(
    dataframe=train_df,
    x_col='Filepath',
    y_col='TS',
    target_size=(256, 256),
    color_mode='grayscale',
    class_mode='raw',
    batch_size=1,
    shuffle=True,
    seed=42,
    subset='validation'
)

test_input = test_generator.flow_from_dataframe(
    dataframe=test_df,
    x_col='Filepath',
    y_col='TS',
    target_size=(256, 256),
    color_mode='grayscale',
    class_mode='raw',
    batch_size=1,
    shuffle=False
)

inputs = tf.keras.Input(shape=(256, 256, 1))
x = tf.keras.layers.Conv2D(filters=32, kernel_size=(3, 3), activation='relu')(inputs)
x = tf.keras.layers.Conv2D(filters=32, kernel_size=(3, 3), activation='relu')(x)
x = tf.keras.layers.MaxPool2D()(x)
x = tf.keras.layers.Conv2D(filters=64, kernel_size=(3, 3), activation='relu')(x)
x = tf.keras.layers.Conv2D(filters=64, kernel_size=(3, 3), activation='relu')(x)
x = tf.keras.layers.MaxPool2D()(x)
x = tf.keras.layers.Conv2D(filters=128, kernel_size=(3, 3), activation='relu')(x)
x = tf.keras.layers.Conv2D(filters=128, kernel_size=(3, 3), activation='relu')(x)
x = tf.keras.layers.MaxPool2D()(x)
x = tf.keras.layers.Flatten()(x)
x = tf.keras.layers.Dense(128, kernel_initializer='he_normal')(x)
x = tf.keras.layers.Dense(64, kernel_initializer='he_normal')(x)

outputs = tf.keras.layers.Dense(1, activation='linear')(x)

model = tf.keras.Model(inputs=inputs, outputs=outputs)

model.compile(
    optimizer='adam',
    loss='mae'
)

history = model.fit(
    train_input,
    validation_data=val_input,
    epochs=10,
    callbacks=[
        tf.keras.callbacks.EarlyStopping(
            monitor='val_loss',
            patience=5,
            restore_best_weights=True
        )
    ]
)

#Results
predicted_TS = np.squeeze(model.predict(test_input))
true_TS = test_input.labels

rmse = np.sqrt(model.evaluate(test_input, verbose=0))
print("     Test RMSE: {:.5f}".format(rmse))

r2 = r2_score(true_TS, predicted_TS)
print("Test R^2 Score: {:.5f}".format(r2))

null_rmse = np.sqrt(np.sum((true_TS - np.mean(true_TS))**2) / len(true_TS))
print("Null/Baseline Model Test RMSE: {:.5f}".format(null_rmse))

预测结果

预测结果示意图

待解答问题

  1. 导致预测值高度趋同的原因是什么?
  2. 可通过哪些手段降低模型的RMSE指标?

问题解答

预测值高度趋同的核心原因

  • 训练配置不合理:仅设置10轮训练远不足以让图像回归模型收敛,同时batch_size设为1会导致梯度波动极大,模型无法学到有效特征映射,只能输出接近标签均值的结果来最小化损失。
  • 数据预处理缺失:未对TS目标值做归一化/标准化处理,抗拉强度数值范围较大,会导致输出层权重更新不稳定,模型最终输出平均化结果。
  • 模型结构缺陷:Flatten后的全连接层未添加ReLU激活函数,特征拟合能力被限制,容易出现梯度消失问题,同时无正则化策略,模型泛化能力差。
  • 数据增强不足:仅对图像做了归一化处理,没有引入随机翻转、旋转、对比度调整等增强操作,模型无法学到鲁棒的纹理特征,极易出现欠拟合。

降低RMSE的优化方案

  • 优化目标值处理:使用StandardScaler或MinMaxScaler对TS标签做缩放,预测完成后再反转换为原始数值,大幅提升模型收敛效率。
  • 调整训练参数:将batch_size调整为832(根据显存容量适配),训练轮次上调到50100轮,配合ReduceLROnPlateau回调动态降低学习率,避免训练提前停滞。
  • 补充数据增强:在训练用ImageDataGenerator中加入随机水平/垂直翻转、小角度旋转、亮度/对比度调整等操作,扩充训练数据分布,提升模型泛化能力。
  • 完善模型结构:在两个全连接层后添加ReLU激活函数,加入Dropout层(dropout比例0.2~0.5)防止过拟合,也可替换为预训练的EfficientNet等骨干网络做迁移学习,提升特征提取能力。
  • 调整优化策略:可尝试使用MSE或MSE+MAE组合损失函数替代单一MAE损失,更适配回归任务的梯度优化,训练过程中监控训练集和验证集损失变化,及时调整训练策略。
  • 校验数据标注:确认图像和TS标签的对应关系正确,避免标注错误导致模型无法学到有效映射。

内容的提问来源于stack exchange,提问作者Gina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 08:06:03