使用Numpy实现线性回归梯度下降时遇维度不匹配报错求助
解决Numpy实现线性回归梯度下降的报错问题
错误原因分析
- 参数维度初始化错误:
- GD函数里,
theta = np.zeros((X.shape[0], X.shape[1]))创建了(1000,5)的数组,但线性回归的参数theta维度应与特征数一致(即(5,)),错误维度导致后续矩阵运算形状不匹配,触发广播错误。 - SGD函数里,
np.zeros(1, X.shape[1])语法错误,正确写法应为np.zeros(X.shape[1])或np.zeros((1, X.shape[1]))。
- GD函数里,
- 梯度计算逻辑错误:
- GD中矩阵乘法顺序错误,预测值应是
X @ theta而非np.dot(theta.T, X);梯度公式应为2 * X.T @ (X @ theta - y) / 样本数,原代码的均值使用和运算顺序均不符合MSE的导数推导。 - SGD中单个样本的梯度无需取均值,直接基于单样本误差计算即可。
- GD中矩阵乘法顺序错误,预测值应是
- MSE函数的维度适配问题:当theta为(5,)时,
X @ theta即可得到与y匹配的(1000,)预测值,无需转置。
修正后的完整代码
import numpy as np # 生成数据集 rng = np.random.RandomState(10) X = 10*rng.rand(1000, 5) # 1000个样本,5个特征 # 目标向量:包含截距0.9,真实特征权重为[2.2,4,-4,1,2] y = 0.9 + np.dot(X, [2.2, 4, -4, 1, 2]) # 批量梯度下降(BGD)实现 def GD(X, y, eta=0.001, n_iter=1000): theta = np.zeros(X.shape[1]) # 初始化参数:维度与特征数一致 m = X.shape[0] # 样本总数 for _ in range(n_iter): error = X @ theta - y # 计算预测误差 grad = 2 * X.T @ error / m # 计算梯度 theta = theta - eta * grad # 更新参数 return theta # 随机梯度下降(SGD)实现 def SGD(X, y, eta=0.001, n_iter=10): theta = np.zeros(X.shape[1]) m = X.shape[0] for _ in range(n_iter): # 打乱样本顺序,避免陷入局部最优 idx = np.random.permutation(m) X_shuffled, y_shuffled = X[idx], y[idx] for j in range(m): error = X_shuffled[j] @ theta - y_shuffled[j] grad = 2 * error * X_shuffled[j] # 单样本梯度 theta = theta - eta * grad return theta # MSE损失函数 def MSE(X, y, theta): return np.mean((X @ theta - y)**2) # 训练并输出结果 theta_gd = GD(X, y) theta_sgd = SGD(X, y) print('真实特征权重:[2.2, 4, -4, 1, 2](注:X未添加偏置项,故theta不含截距0.9)') print('GD训练得到的theta:', theta_gd) print('MSE with GD: ', MSE(X, y, theta_gd)) print('SGD训练得到的theta:', theta_sgd) print('MSE with SGD: ', MSE(X, y, theta_sgd))
额外说明
- 若需要让模型学习截距项,可给X添加一列全为1的偏置特征,此时theta维度会变为6(5个特征权重+1个截距)。
- 学习率
eta设置为0.001是因为X取值范围为0-10,过大的学习率会导致参数震荡无法收敛,可根据实际情况调整。 - SGD中添加样本打乱步骤,能避免固定样本顺序带来的收敛偏差。
内容的提问来源于stack exchange,提问作者Novinnam
相关产品推荐
相关产品推荐

