逻辑回归模型测试得分远低于训练得分,拟合问题求助
模型过拟合与泛化能力差问题的排查与解决
问题背景
拟合Logistic Regression模型时出现严重过拟合:网格搜索得到的最优参数在训练交叉验证中得分达到0.755,但测试集准确率仅0.49,混淆矩阵表现糟糕,且其他模型也存在同类问题。数据集共约900行,按30%比例拆分测试集,采用分层抽样。
原始代码
#importing libraries from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler from sklearn.decomposition import PCA from sklearn.model_selection import GridSearchCV from sklearn.preprocessing import OrdinalEncoder from sklearn.preprocessing import LabelEncoder from sklearn.preprocessing import MinMaxScaler #creating test train splits print(df['high_traffic'].value_counts()) y=df['high_traffic'] X= df.drop(["recipe","high_traffic"],axis=1) X= pd.get_dummies(X,columns=['category']) le = LabelEncoder() X["servings"] = le.fit_transform(X["servings"]) X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=3,stratify=y) #scaling variables scaler = StandardScaler() scaled_train_X = scaler.fit_transform(X_train) #PCA pca=PCA() pca.fit(scaled_train_X) exp_variance = pca.explained_variance_ratio_ cum_exp_variance = np.cumsum(exp_variance) print(cum_exp_variance) pca = PCA(n_components=12,random_state=7) train_pca = pca.fit_transform(scaled_train_X) test_pca = pca.fit_transform(X_test) #logistic regression from sklearn.linear_model import LogisticRegression logreg = LogisticRegression(random_state=3) parameters = [ {'penalty' : ['l1', 'l2', 'elasticnet', 'none'], 'C' : np.logspace(-4, 4, 20), 'solver' : ['lbfgs','newton-cg','liblinear','sag','saga'], 'max_iter' : [100, 1000,2500, 5000] }] grid_search = GridSearchCV(estimator = logreg, param_grid = parameters, scoring = 'accuracy', cv = 5, verbose=1, refit=True) grid_search.fit(train_pca, y_train) #logreg.fit(train_pca,y_train) #pred_y_log = logreg.predict(test_pca) pred_y_log = grid_search.best_estimator_.predict(test_pca) print(grid_search.best_params_) print(grid_search.best_score_) from sklearn.metrics import classification_report class_log = classification_report(y_test,pred_y_log) print("Logistic Regression: \n", class_log) from sklearn.metrics import confusion_matrix print(confusion_matrix(y_test,pred_y_log))
原始运行结果
Fitting 5 folds for each of 1600 candidates, totalling 8000 fits
{'C': 0.23357214690901212, 'max_iter': 100, 'penalty': 'l1', 'solver': 'saga'}
0.7551049197672368
Logistic Regression:
accuracy 0.49
[[54 37]
[80 57]]
核心问题修复与优化方案
1. 修复数据预处理的致命错误
测试集未遵循"训练集拟合、测试集转换"的原则,导致分布偏移:
- 错误代码:
test_pca = pca.fit_transform(X_test) - 修正代码:
# 用训练集拟合的scaler标准化测试集 scaled_test_X = scaler.transform(X_test) # 用训练集拟合的PCA转换测试集 test_pca = pca.transform(scaled_test_X)
2. 特征处理合理性调整
servings字段用LabelEncoder不合理:若该字段是数值型(如1人份、2人份),LabelEncoder会破坏数值的有序关系,建议改用MinMaxScaler或直接保留原始数值。- 检查类别分布:从混淆矩阵看测试集负类134条、正类94条,虽用了分层抽样,但仍需确认原始数据集是否存在更严重的类别不平衡。
3. 模型与评估逻辑优化
- 替换评估指标:样本不平衡时准确率参考性差,网格搜索的
scoring改为f1或roc_auc,模型评估改用F1-score、召回率/精确率。 - 简化模型复杂度:当前网格搜索包含1600个参数组合,小数据集易过拟合,建议先固定参数组合(如
l1对应saga/liblinear,l2对应lbfgs),减少候选数量。 - 增强正则化:C值是正则化强度的倒数,当前最优C=0.23,可尝试更小的C值强化正则化,抑制过拟合。
4. 数据集与验证策略升级
- 小数据集(900行)拆分30%测试集后仅270条样本,结果波动大,建议改用10折交叉验证评估泛化能力,或用
ShuffleSplit多次拆分取平均结果。 - 补充特征工程:尝试组合现有特征、生成统计特征,提升模型的有效学习维度。
内容的提问来源于stack exchange,提问作者Karolis Hematogenas
相关产品推荐
相关产品推荐

