Tensorflow/Keras DNN回归模型精度与模型问题求助
问题描述
我是Tensorflow/Keras新手,用自有数据搭建了首个DNN回归模型,用400个特征预测1个标签,但遇到损失与精度相关问题。多次试验发现,训练轮次的损失和验证精度很早就趋于平稳,验证精度表现不佳。试过100到3000个不同的训练轮次,结果一致,怀疑是代码有bug、模型设置不合理或是训练数据不足。
版本信息
- Tensorflow 2.9.1
- Python 3.9.12
- Numpy 1.23.1
- Pandas 1.4.3
数据说明
原始数据样本:特征为0-400列,标签为400-599列,本模型仅预测标签“lower_pos_0”(第400列)。
代码
import numpy as np import pandas as pd import tensorflow as tf from tensorflow import keras from tensorflow.keras import layers from tensorflow.keras.callbacks import TensorBoard import datetime print(tf.__version__) # Make NumPy printouts easier to read. np.set_printoptions(precision=3, suppress=True) #read in the csv file into a dataframe. Sample file path "C:\...\...\...\csv name.csv" filelocation = "C:\(place the path to the csv here" raw_dataset = pd.read_csv(filelocation) #drop all columns that this model will NOT be used in training this model. Those columns are "lower_pos_1" through "lower_load_99" raw_dataset = raw_dataset.drop(raw_dataset.loc[:, 'lower_pos_1':'lower_load_99'].columns, axis=1) print('shape of raw_dataset:', raw_dataset.shape) print(raw_dataset.head()) #convert the data frame to array dataset = raw_dataset.copy() dataset.tail() #create a training and test set train_dataset = dataset.sample(frac=0.8, random_state=0) test_dataset = dataset.drop(train_dataset.index) #check the data train_dataset.describe().transpose() #split features from labels train_features = train_dataset.copy() test_features = test_dataset.copy() #print the shape print('train features shape:',train_features.shape) print(train_features.head()) print('test features shape:',test_features.shape) print(test_features.head()) #drop the labels from the features. The "lower_pos_0" column is the label that the model is trying to predict. train_labels = train_features.pop('lower_pos_0') test_labels = test_features.pop('lower_pos_0') #print the shape print('train labels shape:',train_labels.shape) print(train_labels.head()) print('test labels shape:',test_labels.shape) print(test_labels.head()) #normalize the data using keras normalizer = tf.keras.layers.Normalization(axis=-1) normalizer.adapt(np.array(train_features)) print(normalizer.mean.numpy()) first = np.array(train_features[:1]) with np.printoptions(precision=2, suppress=True): print('First example:', first) print() print('Normalized:', normalizer(first).numpy()) #build the model def build_and_compile_model(norm): model = keras.Sequential([ norm, layers.Dense(401, activation='relu'), layers.Dense(401, activation='relu'), layers.Dense(1) ]) model.compile(loss='mean_absolute_error', optimizer=tf.keras.optimizers.Adam(0.001), metrics='accuracy') return model #display the models summary dnn_model = build_and_compile_model(normalizer) dnn_model.summary() #setup tensorboard for viewing data during training log_folder = "Traininglogs/" + datetime.datetime.now().strftime("%Y%m%d-%H%M%S") callbacks = [TensorBoard(log_dir=log_folder, histogram_freq=1, write_graph=True, write_images=True, update_freq='epoch', profile_batch=2, embeddings_freq=1)] #train the model with model.fit() dnn_model.fit(train_features, train_labels, epochs=125, validation_split=0.2, callbacks=callbacks) #evaluate the model with model.evaluate() loss, mae = dnn_model.evaluate(test_features, test_labels, verbose=2) #saving the model once completed dnn_model.save('dnn_model')
解决方案
1. 核心错误:回归任务误用分类指标
你的模型是回归任务(预测连续值),但编译时指定了metrics='accuracy'——这个指标仅适用于分类任务,对回归完全无效。这是你看到“验证精度表现不佳”的根本原因,Keras会强行把回归输出和标签做分类精度计算,结果没有任何参考意义。
修改编译代码,换成回归任务适用的指标:
model.compile(loss='mean_absolute_error', optimizer=tf.keras.optimizers.Adam(0.001), metrics=['mean_absolute_error'])
同时调整评估代码的变量名,避免混淆:
loss, val_mae = dnn_model.evaluate(test_features, test_labels, verbose=2)
2. 模型结构优化
- 当前两层401神经元的结构和输入特征数(400)几乎一致,容易引发过拟合或梯度消失。建议缩小层规模,或加入Dropout层抑制过拟合:
其中model = keras.Sequential([ norm, layers.Dense(256, activation='relu', kernel_initializer='he_normal'), layers.Dropout(0.2), layers.Dense(128, activation='relu', kernel_initializer='he_normal'), layers.Dropout(0.2), layers.Dense(1) ])he_normal初始化器配合ReLU激活,能有效缓解梯度消失问题。
3. 训练策略调整
- 添加早停(Early Stopping)回调,自动在验证损失不再下降时停止训练,避免无意义迭代,同时保留最优权重:
from tensorflow.keras.callbacks import EarlyStopping callbacks = [ TensorBoard(log_dir=log_folder, histogram_freq=1, write_graph=True, write_images=True, update_freq='epoch', profile_batch=2, embeddings_freq=1), EarlyStopping(monitor='val_loss', patience=10, restore_best_weights=True) ] - 尝试调整Adam学习率,比如从
0.001降到0.0005,若训练损失下降过慢再适当调高。
4. 数据相关检查
- 确认训练样本量:如果样本数远小于特征数的10倍(比如不足4000条),模型很难学到有效规律,优先补充数据。
- 计算特征与标签的相关性:用
train_features.corrwith(train_labels)输出每个特征和标签的相关系数,剔除相关性极低的特征,减少噪声干扰。 - 检查标签分布:若标签值集中在狭窄区间,可尝试对标签做标准化或对数变换,降低模型学习难度。
5. 代码细节修正
- 文件路径存在语法错误:
"C:\(place the path to the csv here"需改为双反斜杠格式"C:\\path\\to\\your\\file.csv"或原始字符串r"C:\path\to\your\file.csv",否则会触发路径解析错误。 - 移除无意义代码:
dataset = raw_dataset.copy(); dataset.tail()这行没有实际作用,可直接删除。
内容的提问来源于stack exchange,提问作者TheNewGuy
相关产品推荐
相关产品推荐

