手写Linear Regression适配CSV文件时遇IndexError问题求助(梯度下降)
手写线性回归适配CSV数据时的IndexError排查与解决
问题场景
跟随YouTube视频实现基于梯度下降的手写线性回归,原代码使用随机生成的二维数组X、y初始化模型,现在改为读取student-mat.csv的G1列作为X、G2列作为y,运行时触发如下错误:
Traceback (most recent call last): File "/Users/brasilgu/PycharmProjects/LinReg2/main.py", line 51, in <module> w = regression.main(X, y) File "/Users/brasilgu/PycharmProjects/LinReg2/main.py", line 23, in main x1 = np.ones((1, X.shape[1])) IndexError: tuple index out of range
原代码如下:
import numpy as np import pandas as pd class LinearRegression(): def __init__(self): self.learning_rate = 0.001 self.total_iterations = 10000 def y_hat(self, X, w): return np.dot(w.T, X) def loss(self, yhat, y): L =1/self.m * np.sum(np.power(yhat-y, 2)) return L def gradient_descent(self, w, X, y, yhat): dldW = np.dot(X, (yhat - y).T) w = w - self.learning_rate * dldW return w def main(self, X, y): x1 = np.ones((1, X.shape[1])) x = np.append(X, x1, axis=0) self.m = X.shape[1] self.n = X.shape[0] w = np.zeros((self.n, 1)) for it in range(self.total_iterations+1): yhat = self.y_hat(X, w) loss = self.loss(yhat, y) if it % 2000 == 0: print(f'Cost at iteration {it} is {loss}') w = self.gradient_descent(w, X, y, yhat) return w if __name__ == '__main__': #X = np.random.rand(1, 500) #y = 3 * X + np.random.randn(1, 500) * 0.1 data = pd.read_csv('/Users/brasilgu/Downloads/student (1) 2/student-mat.csv', sep=";") X = data['G1'].values y = data['G2'].values regression = LinearRegression() w = regression.main(X, y)
错误原因
- 数组维度不匹配:原代码中随机生成的X是二维数组(形状
(1, 500)),因此X.shape包含两个维度;而从CSV读取的data['G1'].values是一维数组(形状(n,),n为样本数量),仅一个维度,访问X.shape[1]必然触发索引越界。 - 矩阵运算逻辑未适配:原代码的运算基于「特征行、样本列」的格式(特征数×样本数),但读取的一维数组未转换为该格式,后续的矩阵点积、梯度计算都会出现维度不兼容问题。
- 截距项拼接后未使用:原代码中拼接了全1的截距项到
x变量,但后续运算仍用原始X,导致截距项未生效。
解决方法
- 将读取的X、y转换为二维数组,保持「1行特征,n列样本」的格式;
- 修正main方法中使用拼接截距项后的数组参与运算;
- 调整梯度下降和预测函数的维度匹配逻辑。
修改后的完整代码
import numpy as np import pandas as pd class LinearRegression(): def __init__(self): self.learning_rate = 0.001 self.total_iterations = 10000 def y_hat(self, X, w): return np.dot(w.T, X) def loss(self, yhat, y): L = 1/self.m * np.sum(np.power(yhat - y, 2)) return L def gradient_descent(self, w, X, y, yhat): dldW = 2/self.m * np.dot(X, (yhat - y).T) w = w - self.learning_rate * dldW return w def main(self, X, y): # 拼接截距项的全1行,此时X的形状是(2, m):原特征行+截距行 x1 = np.ones((1, X.shape[1])) X = np.append(X, x1, axis=0) self.m = X.shape[1] # 样本数量 self.n = X.shape[0] # 特征数(含截距) w = np.zeros((self.n, 1)) # 参数权重,形状(2,1) for it in range(self.total_iterations + 1): yhat = self.y_hat(X, w) loss = self.loss(yhat, y) if it % 2000 == 0: print(f'Cost at iteration {it} is {loss:.4f}') w = self.gradient_descent(w, X, y, yhat) return w if __name__ == '__main__': data = pd.read_csv('/Users/brasilgu/Downloads/student (1) 2/student-mat.csv', sep=";") # 将一维数组转换为二维数组:(1, 样本数) X = data['G1'].values.reshape(1, -1) y = data['G2'].values.reshape(1, -1) regression = LinearRegression() w = regression.main(X, y) print(f"训练得到的权重:{w}")
额外说明
- 梯度下降公式中补充了
2/self.m的系数,这是均方误差损失的导数正确形式; - 输出损失时保留4位小数,便于观察收敛情况;
- 最后打印训练得到的权重,方便验证结果。
内容的提问来源于stack exchange,提问作者philomath
相关产品推荐
相关产品推荐

