为何添加冗余特征会提升Logistic Regression模型的准确率?
为什么添加冗余特征后Logistic回归的准确率会上升?
Logistic回归实验代码
from sklearn.datasets import load_iris from sklearn.linear_model import LogisticRegression import numpy as np X, y = load_iris(return_X_y=True) for i in range(5): X_redundant = np.c_[X,X[:,:i]] # 重复添加冗余特征 print(X_redundant.shape) clf = LogisticRegression(random_state=0,max_iter=1000).fit(X_redundant, y) print(clf.score(X_redundant, y))
实验输出结果
(150, 4) 0.9733333333333334 (150, 5) 0.98 (150, 6) 0.98 (150, 7) 0.9866666666666667 (150, 8) 0.9866666666666667
问题
为什么添加冗余特征时,Logistic Regression的分数(默认指标为准确率)会上升?我原本预期分数会保持不变,这是类比LinearRegression的行为得出的结论。
线性回归对比实验
对于LinearRegression而言,添加更多特征列后分数(默认指标为R²)不会改变,因为LinearRegression会将系数均匀分配给每一对冗余特征。
线性回归实验代码
from sklearn.datasets import load_iris from sklearn.linear_model import LinearRegression import numpy as np X, y = load_iris(return_X_y=True) X, y = X[:,:-1],X[:,-1] for i in range(4): X_redundant = np.c_[X,X[:,:i]] # 重复添加冗余特征 print(X_redundant.shape) clf = LinearRegression().fit(X_redundant, y) print(clf.score(X_redundant, y)) print(clf.coef_)
线性回归实验输出结果
(150, 3) 0.9378502736046809 [-0.20726607 0.22282854 0.52408311] (150, 4) 0.9378502736046809 [-0.10363304 0.22282854 0.52408311 -0.10363304] (150, 5) 0.9378502736046809 [-0.10363304 0.11141427 0.52408311 -0.10363304 0.11141427] (150, 6) 0.9378502736046809 [-0.10363304 0.11141427 0.26204156 -0.10363304 0.11141427 0.26204156]
原因解析
核心差异在于Logistic回归默认带有L2正则化,而sklearn中的线性回归默认无正则化:
线性回归的表现逻辑
线性回归最小化平方损失,当加入冗余特征时,模型会将原特征的权重拆分到冗余特征上(比如原权重w拆成两个各为w/2的权重),整体损失值完全不变,因此R²分数保持一致,不受冗余特征影响。Logistic回归的表现逻辑
sklearn的LogisticRegression默认启用penalty='l2'(L2正则化)和C=1.0(正则化强度的倒数)。L2正则化会惩罚所有权重的平方和:
- 加入冗余特征后,原特征x的权重w会被拆分为多个冗余特征的权重(比如w1=w2=w/2),此时权重平方和从w²变为(w/2)²+(w/2)²=w²/2,正则化项的惩罚力度直接减半。
- 正则化强度降低意味着模型对训练数据的拟合约束减少,原本处于分类边界的样本更容易被正确分类,因此训练集上的准确率会上升。
如果验证这个结论,可以把Logistic回归的正则化关闭(设置penalty='none'),此时添加冗余特征后准确率将不再变化,和线性回归表现一致。
内容的提问来源于stack exchange,提问作者Han Qi
相关产品推荐
相关产品推荐

