You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python手动实现线性回归:数组值未更新及NaN问题排查

线性回归实现NaN问题与参数不更新排查

嘿,我帮你排查下这段手动实现线性回归代码里的问题——你遇到的NaN输出和参数“未更新”(实际是参数直接爆炸到NaN)主要有几个核心问题:

1. 损失函数计算逻辑错误

你写的costfn里,np.sum(x.dot(theta.T) - y) ** 2是先把所有样本的预测误差求和,再整体平方,这完全偏离了线性回归的均方误差(MSE)损失逻辑!

正确的计算应该是先对每个样本的误差单独平方,再求和,也就是:

def costfn(x, y, theta):
    j = np.sum((x.dot(theta.T) - y.reshape(-1,1)) ** 2) / (2 * len(y))
    return j

原写法会让损失值计算完全失真,直接导致梯度更新的方向和幅度出错。

2. 特征未标准化,梯度爆炸触发NaN

波士顿房价数据集的特征尺度差异极大:比如TAX特征值可达数百,而CRIM是0到几十的小数。这种情况下用alpha=0.001的学习率会让梯度更新步长过大,参数瞬间溢出变成无穷大,最终计算时输出NaN。

解决方法是对训练集特征做标准化(均值为0,方差为1),测试集也要用训练集的均值和方差做同样处理:

# 在数据集拆分后添加标准化步骤
mean = np.mean(x_train, axis=0)
std = np.std(x_train, axis=0)
x_train = (x_train - mean) / std
x_test = (x_test - mean) / std

3. 矩阵运算维度可优化(非致命但建议调整)

你的假设函数h = theta.dot(x.T)虽然维度合法,但和costfn里的x.dot(theta.T)计算逻辑不一致,容易混淆。另外,把一维的y_train转成二维数组,能避免广播运算的潜在问题:

def gradient(x, y, theta, alpha, iterations):
    cost_history = [0] * iterations
    y = y.reshape(-1,1)  # 将y转为(N,1)的二维数组
    for i in range(iterations):
        h = x.dot(theta.T)  # 和costfn保持一致的计算逻辑
        loss = h - y
        g = loss.T.dot(x) / len(y)  # 转置loss后与x相乘,得到(1,13)的梯度
        theta = theta - alpha * g
        cost_history[i] = costfn(x, y, theta)
    return theta, cost_history

修正后的完整代码

同时注意sklearn.cross_validation已被废弃,改用sklearn.model_selection:

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split

data = pd.read_csv("housing.csv", delimiter=' ', skipinitialspace=True, names=['CRIM', 'ZN', 'INDUS', 'CHAS', 'NOX', 'RM', 'AGE', 'DIS', 'RAD', 'TAX', 'PTRATIO', 'B', 'LSTAT', 'MEDV'])
df_x = data.drop('MEDV', axis=1)
df_y = data['MEDV']
x_train, x_test, y_train, y_test = train_test_split(df_x.values, df_y.values, test_size=0.2, random_state=4)

# 特征标准化
mean = np.mean(x_train, axis=0)
std = np.std(x_train, axis=0)
x_train = (x_train - mean) / std
x_test = (x_test - mean) / std

theta = np.zeros((1, 13))

def costfn(x, y, theta):
    j = np.sum((x.dot(theta.T) - y.reshape(-1,1)) ** 2) / (2 * len(y))
    return j

def gradient(x, y, theta, alpha, iterations):
    cost_history = [0] * iterations
    y = y.reshape(-1,1)
    for i in range(iterations):
        h = x.dot(theta.T)
        loss = h - y
        g = loss.T.dot(x) / len(y)
        theta = theta - alpha * g
        cost_history[i] = costfn(x, y, theta)
    return theta, cost_history

# 标准化后可适当提高学习率
theta, cost_history = gradient(x_train, y_train, theta, 0.01, 1000)
print(theta)
print(f"最终损失值: {cost_history[-1]}")

内容的提问来源于stack exchange,提问作者parth shukla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:05:59