Python手动实现线性回归:数组值未更新及NaN问题排查
线性回归实现NaN问题与参数不更新排查
嘿,我帮你排查下这段手动实现线性回归代码里的问题——你遇到的NaN输出和参数“未更新”(实际是参数直接爆炸到NaN)主要有几个核心问题:
1. 损失函数计算逻辑错误
你写的costfn里,np.sum(x.dot(theta.T) - y) ** 2是先把所有样本的预测误差求和,再整体平方,这完全偏离了线性回归的均方误差(MSE)损失逻辑!
正确的计算应该是先对每个样本的误差单独平方,再求和,也就是:
def costfn(x, y, theta): j = np.sum((x.dot(theta.T) - y.reshape(-1,1)) ** 2) / (2 * len(y)) return j
原写法会让损失值计算完全失真,直接导致梯度更新的方向和幅度出错。
2. 特征未标准化,梯度爆炸触发NaN
波士顿房价数据集的特征尺度差异极大:比如TAX特征值可达数百,而CRIM是0到几十的小数。这种情况下用alpha=0.001的学习率会让梯度更新步长过大,参数瞬间溢出变成无穷大,最终计算时输出NaN。
解决方法是对训练集特征做标准化(均值为0,方差为1),测试集也要用训练集的均值和方差做同样处理:
# 在数据集拆分后添加标准化步骤 mean = np.mean(x_train, axis=0) std = np.std(x_train, axis=0) x_train = (x_train - mean) / std x_test = (x_test - mean) / std
3. 矩阵运算维度可优化(非致命但建议调整)
你的假设函数h = theta.dot(x.T)虽然维度合法,但和costfn里的x.dot(theta.T)计算逻辑不一致,容易混淆。另外,把一维的y_train转成二维数组,能避免广播运算的潜在问题:
def gradient(x, y, theta, alpha, iterations): cost_history = [0] * iterations y = y.reshape(-1,1) # 将y转为(N,1)的二维数组 for i in range(iterations): h = x.dot(theta.T) # 和costfn保持一致的计算逻辑 loss = h - y g = loss.T.dot(x) / len(y) # 转置loss后与x相乘,得到(1,13)的梯度 theta = theta - alpha * g cost_history[i] = costfn(x, y, theta) return theta, cost_history
修正后的完整代码
同时注意sklearn.cross_validation已被废弃,改用sklearn.model_selection:
import pandas as pd import numpy as np import matplotlib.pyplot as plt from sklearn.model_selection import train_test_split data = pd.read_csv("housing.csv", delimiter=' ', skipinitialspace=True, names=['CRIM', 'ZN', 'INDUS', 'CHAS', 'NOX', 'RM', 'AGE', 'DIS', 'RAD', 'TAX', 'PTRATIO', 'B', 'LSTAT', 'MEDV']) df_x = data.drop('MEDV', axis=1) df_y = data['MEDV'] x_train, x_test, y_train, y_test = train_test_split(df_x.values, df_y.values, test_size=0.2, random_state=4) # 特征标准化 mean = np.mean(x_train, axis=0) std = np.std(x_train, axis=0) x_train = (x_train - mean) / std x_test = (x_test - mean) / std theta = np.zeros((1, 13)) def costfn(x, y, theta): j = np.sum((x.dot(theta.T) - y.reshape(-1,1)) ** 2) / (2 * len(y)) return j def gradient(x, y, theta, alpha, iterations): cost_history = [0] * iterations y = y.reshape(-1,1) for i in range(iterations): h = x.dot(theta.T) loss = h - y g = loss.T.dot(x) / len(y) theta = theta - alpha * g cost_history[i] = costfn(x, y, theta) return theta, cost_history # 标准化后可适当提高学习率 theta, cost_history = gradient(x_train, y_train, theta, 0.01, 1000) print(theta) print(f"最终损失值: {cost_history[-1]}")
内容的提问来源于stack exchange,提问作者parth shukla
相关产品推荐
相关产品推荐

