线性回归模型测试阶段RMSE过高问题的排查与解决
自定义线性回归模型预测房价异常问题解决方案
问题概述
自定义实现的LinearRegression模型训练时MSE处于合理范围,但测试集RMSE极高,甚至出现负房价预测值,评估准确率为0。数据集包含14个特征和1个房价标签。
问题复现代码
import numpy as np import pandas as pd pd.set_option('future.no_silent_downcasting', True) class LinearRegression: def __init__(self, x_train, y_train, epochs=20, alpha=0.01): self.x_train = pd.DataFrame(x_train) self.y_train = pd.DataFrame(y_train).values.reshape(-1, 1) self.shape = x_train.shape self.x_scale_factor = self.x_train.max(axis=0) self.y_scale_factor = self.y_train.max() self.x_train /= self.x_scale_factor self.y_train /= self.y_scale_factor self.weight_matrix = np.random.rand(1, self.shape[1]) self.bias_matrix = np.random.rand(1, 1) self.epochs = epochs self.alpha = alpha self.error = 0 def train(self): for i in range(self.epochs): predictions = np.dot(self.x_train, self.weight_matrix.T) + self.bias_matrix error_matrix = ((self.y_train - predictions) ** 2) / self.shape[0] self.error = np.sum(error_matrix) weight_gradient = -2 * np.dot((self.y_train - predictions).T, self.x_train) / self.shape[0] bias_gradient = -2 * np.mean(self.y_train - predictions) self.weight_matrix -= self.alpha * weight_gradient self.bias_matrix -= self.alpha * bias_gradient if i % 10 == 0: print(f"Epoch {i} \t error: {self.error}") print(self.weight_matrix) print(self.bias_matrix) def predict(self, x_features): x_features = pd.DataFrame(x_features) x_features /= self.x_scale_factor y_predictions = np.dot(x_features, self.weight_matrix.T) + self.bias_matrix y_predictions *= self.y_scale_factor return y_predictions def evaluate(self, x_test, y_test): x_test, y_test = pd.DataFrame(x_test), pd.DataFrame(y_test).values.reshape(-1, 1) x_test /= self.x_scale_factor y_predict = self.predict(x_test) rmse = np.sqrt(np.mean((y_predict - y_test) ** 2)) return rmse # Example usage train_data = pd.read_csv('Dataset/House price/df_train.csv') train_data = train_data.drop('date', axis=1) x_train = train_data.drop(columns=['price']).sample(frac=1) x_train = x_train.replace({True: 1, False: 0}).astype(int) y_train = train_data['price'].sample(frac=1) model = LinearRegression(x_train, y_train, epochs=500, alpha=0.1) model.train() test_data = pd.read_csv('Dataset/House price/df_test.csv') x_test = test_data.drop(columns=['price', 'date']) x_test = x_test.replace({True: 1, False: 0}).astype(int) y_test = test_data['price'] print(f"Root Mean Squared Error: {model.evaluate(x_test, y_test)}")
问题分析与解决步骤
1. 修复特征重复缩放的致命错误
原evaluate方法中手动对测试集做了一次缩放,调用predict时又做了一次缩放,导致特征被过度缩放,直接引发预测值异常。修改evaluate方法:
def evaluate(self, x_test, y_test): y_test = pd.DataFrame(y_test).values.reshape(-1, 1) y_predict = self.predict(x_test) rmse = np.sqrt(np.mean((y_predict - y_test) ** 2)) # 可添加MAE等辅助评估指标 mae = np.mean(np.abs(y_predict - y_test)) print(f"MAE: {mae}") return rmse
2. 替换不稳定的归一化方式
原用最大值做归一化,若测试集存在比训练集更大的特征值,会导致特征缩放后超过1,结合权重计算易出现负数。改用Z-score标准化:
# 修改__init__中的缩放逻辑 def __init__(self, x_train, y_train, epochs=20, alpha=0.01): self.x_train = pd.DataFrame(x_train) self.y_train = pd.DataFrame(y_train).values.reshape(-1, 1) self.shape = x_train.shape # Z-score标准化 self.x_mean = self.x_train.mean(axis=0) self.x_std = self.x_train.std(axis=0) self.x_train = (self.x_train - self.x_mean) / self.x_std self.y_mean = self.y_train.mean() self.y_std = self.y_train.std() self.y_train = (self.y_train - self.y_mean) / self.y_std self.weight_matrix = np.random.normal(0, 0.01, (1, self.shape[1])) # 改用正态分布初始化权重 self.bias_matrix = np.zeros((1,1)) # 偏置初始化为0 self.epochs = epochs self.alpha = alpha self.error = 0 # 修改predict方法中的缩放逻辑 def predict(self, x_features): x_features = pd.DataFrame(x_features) x_features = (x_features - self.x_mean) / self.x_std y_predictions = np.dot(x_features, self.weight_matrix.T) + self.bias_matrix y_predictions = y_predictions * self.y_std + self.y_mean y_predictions = np.maximum(y_predictions, 0) # 截断负预测值为0 return y_predictions
3. 解决过拟合问题
训练集MSE低但测试集RMSE高,属于典型过拟合,可通过以下方式缓解:
- 添加L2正则化,修改
train方法中的损失和梯度计算:def train(self): lambda_reg = 0.01 # 正则化系数,可根据效果调整 for i in range(self.epochs): predictions = np.dot(self.x_train, self.weight_matrix.T) + self.bias_matrix # 加入L2正则项 error_matrix = ((self.y_train - predictions) ** 2) / self.shape[0] + lambda_reg * np.sum(self.weight_matrix **2)/self.shape[0] self.error = np.sum(error_matrix) # 权重梯度加入正则项 weight_gradient = -2 * np.dot((self.y_train - predictions).T, self.x_train) / self.shape[0] + 2 * lambda_reg * self.weight_matrix / self.shape[0] bias_gradient = -2 * np.mean(self.y_train - predictions) self.weight_matrix -= self.alpha * weight_gradient self.bias_matrix -= self.alpha * bias_gradient if i % 10 == 0: print(f"Epoch {i} \t error: {self.error}") - 增加早停机制:训练时监控验证集损失,当连续多轮损失不再下降时提前停止训练。
- 筛选特征:移除与房价无关或相关性极低的特征,减少模型复杂度。
内容的提问来源于stack exchange,提问作者aaditya jindal
相关产品推荐
相关产品推荐

