TensorFlow Keras回归模型拟合数据偏差问题求助
回归模型预测与真实数据偏差过大问题求助
近期我尝试在项目中引入TensorFlow,参照官方《Basic Regression Using Keras Guide》实现回归任务,但遇到模型拟合线与真实数据偏差较大的问题,相关loss图、预测vs数据图已提供。我已对数据进行归一化处理,训练了1000个epoch,数据本身无异常。以下是所用数据链接与代码,恳请各位帮忙分析为何预测结果与真实数据差异显著。
代码实现
import matplotlib.pyplot as plt import numpy as np import pandas as pd import seaborn as sns # Make NumPy printouts easier to read. np.set_printoptions(precision=3, suppress=True) import tensorflow as tf from tensorflow import keras from tensorflow.keras import layers print(tf.__version__) train_dataset = df.sample(frac=0.8, random_state = 0) test_dataset = df.drop(train_dataset.index) train_dataset.describe().transpose() train_features = train_dataset.copy() test_features = test_dataset.copy() train_labels = train_features.pop('Max') test_labels = test_features.pop('Max') train_dataset.describe().transpose()[['mean','std']] normalizer = tf.keras.layers.Normalization(axis=-1) normalizer.adapt(np.array(train_features)) print(normalizer.mean.numpy()) first = np.array(train_features[:1]) with np.printoptions(precision=2, suppress=True): print('First example:', first) print() print('Normalized:', normalizer(first).numpy()) date = np.array(train_features['Date Lifted']) date_normalizer = layers.Normalization(input_shape=[1,], axis=None) date_normalizer.adapt(date) date_model = tf.keras.Sequential([ date_normalizer, layers.Dense(units=1) ]) date_model.summary() date_model.predict(date[:10]) date_model.compile( optimizer=tf.keras.optimizers.Adam(learning_rate=0.001), loss='mean_absolute_error') %%time history = date_model.fit( train_features['Date Lifted'], train_labels, epochs=100, # Suppress logging. verbose=0, # Calculate validation results on 20% of the training data. validation_split = 0.2) hist = pd.DataFrame(history.history) hist['epoch'] = history.epoch hist.tail() def plot_loss(history): plt.plot(history.history['loss'], label='loss') plt.plot(history.history['val_loss'], label='val_loss') plt.ylim([0, 1000]) plt.xlabel('Epoch') plt.ylabel('Error [Max]') plt.legend() plt.grid(True) plot_loss(history) test_results = {} test_results['date_model'] = date_model.evaluate( test_features['Date Lifted'], test_labels, verbose=0) x = tf.linspace(0, 250, 251) y = date_model.predict(x) def plot_horsepower(x, y): plt.scatter(train_features['Date Lifted'], train_labels, label='Data') plt.plot(x, y, color='k', label='Predictions') plt.xlabel('Date Lifted') plt.ylabel('Max') plt.legend() plot_horsepower(x, y)
问题分析与解决建议
- 模型结构过于简单:当前仅使用一个
Dense(units=1)线性层,只能拟合线性关系。若Date Lifted与Max存在非线性关联,模型完全无法捕捉这种规律,必然导致偏差。建议增加模型复杂度,比如添加带激活函数的隐藏层:date_model = tf.keras.Sequential([ date_normalizer, layers.Dense(64, activation='relu'), layers.Dense(64, activation='relu'), layers.Dense(1) ]) - 训练epoch不一致:你提到训练了1000个epoch,但代码中
fit函数设置的是epochs=100,这会导致实际训练轮次不足。若确实训练了1000轮,需确认训练过程中loss是否已趋于稳定;若未达到,需调整代码中的epoch参数。 - 未充分利用特征:当前模型仅使用
Date Lifted单一特征,若数据集存在其他与Max相关的特征,完全丢弃会丢失关键信息。建议尝试使用所有特征构建多特征模型:multi_model = tf.keras.Sequential([ normalizer, layers.Dense(64, activation='relu'), layers.Dense(64, activation='relu'), layers.Dense(1) ]) - 超参数与数据范围检查:
- 学习率:当前
Adam优化器学习率为0.001,若调整模型复杂度,可尝试降低学习率(如0.0001)以提升收敛效果; - 损失函数:可尝试替换为
mean_squared_error,观察拟合效果变化; - 预测范围:确认
train_features['Date Lifted']的实际取值范围是否在0-250之间,若实际数据超出该区间,预测线会与真实数据脱节。
- 学习率:当前
内容的提问来源于stack exchange,提问作者Ashwin Chembu
相关产品推荐
相关产品推荐

