You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow Keras回归模型拟合数据偏差问题求助

回归模型预测与真实数据偏差过大问题求助

近期我尝试在项目中引入TensorFlow,参照官方《Basic Regression Using Keras Guide》实现回归任务,但遇到模型拟合线与真实数据偏差较大的问题,相关loss图、预测vs数据图已提供。我已对数据进行归一化处理,训练了1000个epoch,数据本身无异常。以下是所用数据链接与代码,恳请各位帮忙分析为何预测结果与真实数据差异显著。

代码实现

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import seaborn as sns

# Make NumPy printouts easier to read.
np.set_printoptions(precision=3, suppress=True)
import tensorflow as tf

from tensorflow import keras
from tensorflow.keras import layers

print(tf.__version__)
train_dataset = df.sample(frac=0.8, random_state = 0)
test_dataset = df.drop(train_dataset.index)
train_dataset.describe().transpose()
train_features = train_dataset.copy()
test_features = test_dataset.copy()

train_labels = train_features.pop('Max')
test_labels = test_features.pop('Max')
train_dataset.describe().transpose()[['mean','std']]
normalizer = tf.keras.layers.Normalization(axis=-1)
normalizer.adapt(np.array(train_features))
print(normalizer.mean.numpy())
first = np.array(train_features[:1])

with np.printoptions(precision=2, suppress=True):
  print('First example:', first)
  print()
  print('Normalized:', normalizer(first).numpy())
date = np.array(train_features['Date Lifted'])

date_normalizer = layers.Normalization(input_shape=[1,], axis=None)
date_normalizer.adapt(date)
date_model = tf.keras.Sequential([
    date_normalizer,
    layers.Dense(units=1)
])

date_model.summary()
date_model.predict(date[:10])
date_model.compile(
    optimizer=tf.keras.optimizers.Adam(learning_rate=0.001),
    loss='mean_absolute_error')
%%time
history = date_model.fit(
    train_features['Date Lifted'],
    train_labels,
    epochs=100,
    # Suppress logging.
    verbose=0,
    # Calculate validation results on 20% of the training data.
    validation_split = 0.2)
hist = pd.DataFrame(history.history)
hist['epoch'] = history.epoch
hist.tail()
def plot_loss(history):
  plt.plot(history.history['loss'], label='loss')
  plt.plot(history.history['val_loss'], label='val_loss')
  plt.ylim([0, 1000])
  plt.xlabel('Epoch')
  plt.ylabel('Error [Max]')
  plt.legend()
  plt.grid(True)
plot_loss(history)
test_results = {}

test_results['date_model'] = date_model.evaluate(
    test_features['Date Lifted'],
    test_labels, verbose=0)
x = tf.linspace(0, 250, 251)
y = date_model.predict(x)
def plot_horsepower(x, y):
  plt.scatter(train_features['Date Lifted'], train_labels, label='Data')
  plt.plot(x, y, color='k', label='Predictions')
  plt.xlabel('Date Lifted')
  plt.ylabel('Max')
  plt.legend()
plot_horsepower(x, y)

问题分析与解决建议

  • 模型结构过于简单:当前仅使用一个Dense(units=1)线性层,只能拟合线性关系。若Date Lifted与Max存在非线性关联,模型完全无法捕捉这种规律,必然导致偏差。建议增加模型复杂度,比如添加带激活函数的隐藏层:
    date_model = tf.keras.Sequential([
        date_normalizer,
        layers.Dense(64, activation='relu'),
        layers.Dense(64, activation='relu'),
        layers.Dense(1)
    ])
    
  • 训练epoch不一致:你提到训练了1000个epoch,但代码中fit函数设置的是epochs=100,这会导致实际训练轮次不足。若确实训练了1000轮,需确认训练过程中loss是否已趋于稳定;若未达到,需调整代码中的epoch参数。
  • 未充分利用特征:当前模型仅使用Date Lifted单一特征,若数据集存在其他与Max相关的特征,完全丢弃会丢失关键信息。建议尝试使用所有特征构建多特征模型:
    multi_model = tf.keras.Sequential([
        normalizer,
        layers.Dense(64, activation='relu'),
        layers.Dense(64, activation='relu'),
        layers.Dense(1)
    ])
    
  • 超参数与数据范围检查:
    • 学习率:当前Adam优化器学习率为0.001,若调整模型复杂度,可尝试降低学习率(如0.0001)以提升收敛效果;
    • 损失函数:可尝试替换为mean_squared_error,观察拟合效果变化;
    • 预测范围:确认train_features['Date Lifted']的实际取值范围是否在0-250之间,若实际数据超出该区间,预测线会与真实数据脱节。

内容的提问来源于stack exchange,提问作者Ashwin Chembu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 04:15:34