每次运行线性回归(LR)模型代码得到不同系数是否正常?
线性回归模型每次运行结果不同的原因及解决办法
问题描述
用户编写的线性回归代码每次运行时,自变量系数、截距及MSE值都会变化,无法得到固定回归方程。代码如下:
#REGRESSION ANALYSIS #splitting the dataset into x and y variables firm1=pd.DataFrame(firm, columns=['Sales', 'Advert', 'Empl', 'Prod']) print(firm1) x = firm1.drop(['Sales'], axis=1) y = firm1['Sales'] print(x) print(y) x_train, x_test, y_train, y_test = train_test_split(x,y, test_size=0.2) print(x_train.shape, y_train.shape) print(x_test.shape, y_test.shape) #the LR model M=linear_model.LinearRegression(fit_intercept=True) M.fit(x_train, y_train) y_pred=M.predict(x_test) print(y_pred) print('Coeff: ', M.coef_) for i in M.coef_: print('{:.4f}'.format(i)) print('Intercept: ','{:.4f}'.format(M.intercept_)) print('MSE: ','{:.4f}'.format(mean_squared_error(y_test, y_pred))) print('Coeffieicnt of determination (r2): ','{:.4f}'.format(r2_score(y_test, y_pred))) print(firm1.sample())
第一次运行输出:
Coeff: [454.83981664 63.77031531 59.31844506]
454.8398
63.7703
59.3184
Intercept: -1073.5124
MSE: 434529.9361
第二次运行输出:
Coeff: [462.0304152 61.17909189 269.41075305]
462.0304
61.1791
269.4108
Intercept: -1462.2449
MSE: 4014768.0049
用户咨询这种每次运行结果不同的情况是否正常。
原因分析
这种情况并非正常的预期结果,核心原因是代码中划分训练集和测试集时没有固定随机种子。train_test_split函数默认会随机打乱数据后划分,每次运行时随机数生成器的状态不同,导致划分出的训练集、测试样本不一样。而线性回归模型的参数是基于训练集拟合出来的,训练数据变了,拟合出的系数、截距自然会变;同时测试集不同,计算出的MSE也会跟着变化。
解决办法
在调用train_test_split时,添加random_state参数并设置一个固定值(比如42、100等任意整数),这样每次运行时数据划分的结果就会完全一致,模型训练结果也会固定:
x_train, x_test, y_train, y_test = train_test_split(x,y, test_size=0.2, random_state=42)
设置固定随机种子后,不管运行多少次,训练集和测试集的样本都不会变,拟合出的回归系数、截距以及MSE值都会保持一致。
内容的提问来源于stack exchange,提问作者Arnold Odhiambo
相关产品推荐
相关产品推荐

