You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

线性回归模型测试阶段RMSE过高问题的排查与解决

自定义线性回归模型预测房价异常问题解决方案

问题概述

自定义实现的LinearRegression模型训练时MSE处于合理范围,但测试集RMSE极高,甚至出现负房价预测值,评估准确率为0。数据集包含14个特征和1个房价标签。

问题复现代码

import numpy as np
import pandas as pd

pd.set_option('future.no_silent_downcasting', True)

class LinearRegression:
    def __init__(self, x_train, y_train, epochs=20, alpha=0.01):
        self.x_train = pd.DataFrame(x_train)
        self.y_train = pd.DataFrame(y_train).values.reshape(-1, 1)
        self.shape = x_train.shape
        self.x_scale_factor = self.x_train.max(axis=0)
        self.y_scale_factor = self.y_train.max()
        self.x_train /= self.x_scale_factor
        self.y_train /= self.y_scale_factor
        self.weight_matrix = np.random.rand(1, self.shape[1])
        self.bias_matrix = np.random.rand(1, 1)
        self.epochs = epochs
        self.alpha = alpha
        self.error = 0

    def train(self):
        for i in range(self.epochs):
            predictions = np.dot(self.x_train, self.weight_matrix.T) + self.bias_matrix
            error_matrix = ((self.y_train - predictions) ** 2) / self.shape[0]
            self.error = np.sum(error_matrix)
            weight_gradient = -2 * np.dot((self.y_train - predictions).T, self.x_train) / self.shape[0]
            bias_gradient = -2 * np.mean(self.y_train - predictions)
            self.weight_matrix -= self.alpha * weight_gradient
            self.bias_matrix -= self.alpha * bias_gradient

            if i % 10 == 0:
                print(f"Epoch {i} \t error: {self.error}")
        print(self.weight_matrix)
        print(self.bias_matrix)

    def predict(self, x_features):
        x_features = pd.DataFrame(x_features)
        x_features /= self.x_scale_factor
        y_predictions = np.dot(x_features, self.weight_matrix.T) + self.bias_matrix
        y_predictions *= self.y_scale_factor
        return y_predictions

    def evaluate(self, x_test, y_test):
        x_test, y_test = pd.DataFrame(x_test), pd.DataFrame(y_test).values.reshape(-1, 1)
        x_test /= self.x_scale_factor
        y_predict = self.predict(x_test)
        rmse = np.sqrt(np.mean((y_predict - y_test) ** 2))
        return rmse

# Example usage
train_data = pd.read_csv('Dataset/House price/df_train.csv')
train_data = train_data.drop('date', axis=1)
x_train = train_data.drop(columns=['price']).sample(frac=1)
x_train = x_train.replace({True: 1, False: 0}).astype(int)
y_train = train_data['price'].sample(frac=1)
model = LinearRegression(x_train, y_train, epochs=500, alpha=0.1)
model.train()

test_data = pd.read_csv('Dataset/House price/df_test.csv')
x_test = test_data.drop(columns=['price', 'date'])
x_test = x_test.replace({True: 1, False: 0}).astype(int)
y_test = test_data['price']

print(f"Root Mean Squared Error: {model.evaluate(x_test, y_test)}")

问题分析与解决步骤

1. 修复特征重复缩放的致命错误

原evaluate方法中手动对测试集做了一次缩放,调用predict时又做了一次缩放,导致特征被过度缩放,直接引发预测值异常。修改evaluate方法:

def evaluate(self, x_test, y_test):
    y_test = pd.DataFrame(y_test).values.reshape(-1, 1)
    y_predict = self.predict(x_test)
    rmse = np.sqrt(np.mean((y_predict - y_test) ** 2))
    # 可添加MAE等辅助评估指标
    mae = np.mean(np.abs(y_predict - y_test))
    print(f"MAE: {mae}")
    return rmse

2. 替换不稳定的归一化方式

原用最大值做归一化,若测试集存在比训练集更大的特征值,会导致特征缩放后超过1,结合权重计算易出现负数。改用Z-score标准化:

# 修改__init__中的缩放逻辑
def __init__(self, x_train, y_train, epochs=20, alpha=0.01):
    self.x_train = pd.DataFrame(x_train)
    self.y_train = pd.DataFrame(y_train).values.reshape(-1, 1)
    self.shape = x_train.shape
    # Z-score标准化
    self.x_mean = self.x_train.mean(axis=0)
    self.x_std = self.x_train.std(axis=0)
    self.x_train = (self.x_train - self.x_mean) / self.x_std

    self.y_mean = self.y_train.mean()
    self.y_std = self.y_train.std()
    self.y_train = (self.y_train - self.y_mean) / self.y_std
    
    self.weight_matrix = np.random.normal(0, 0.01, (1, self.shape[1]))  # 改用正态分布初始化权重
    self.bias_matrix = np.zeros((1,1))  # 偏置初始化为0
    self.epochs = epochs
    self.alpha = alpha
    self.error = 0

# 修改predict方法中的缩放逻辑
def predict(self, x_features):
    x_features = pd.DataFrame(x_features)
    x_features = (x_features - self.x_mean) / self.x_std
    y_predictions = np.dot(x_features, self.weight_matrix.T) + self.bias_matrix
    y_predictions = y_predictions * self.y_std + self.y_mean
    y_predictions = np.maximum(y_predictions, 0)  # 截断负预测值为0
    return y_predictions

3. 解决过拟合问题

训练集MSE低但测试集RMSE高,属于典型过拟合,可通过以下方式缓解:

  • 添加L2正则化,修改train方法中的损失和梯度计算:
    def train(self):
        lambda_reg = 0.01  # 正则化系数,可根据效果调整
        for i in range(self.epochs):
            predictions = np.dot(self.x_train, self.weight_matrix.T) + self.bias_matrix
            # 加入L2正则项
            error_matrix = ((self.y_train - predictions) ** 2) / self.shape[0] + lambda_reg * np.sum(self.weight_matrix **2)/self.shape[0]
            self.error = np.sum(error_matrix)
            # 权重梯度加入正则项
            weight_gradient = -2 * np.dot((self.y_train - predictions).T, self.x_train) / self.shape[0] + 2 * lambda_reg * self.weight_matrix / self.shape[0]
            bias_gradient = -2 * np.mean(self.y_train - predictions)
            self.weight_matrix -= self.alpha * weight_gradient
            self.bias_matrix -= self.alpha * bias_gradient
    
            if i % 10 == 0:
                print(f"Epoch {i} \t error: {self.error}")
    
  • 增加早停机制:训练时监控验证集损失,当连续多轮损失不再下降时提前停止训练。
  • 筛选特征:移除与房价无关或相关性极低的特征,减少模型复杂度。

内容的提问来源于stack exchange,提问作者aaditya jindal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 13:52:13