You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

逻辑回归模型测试得分远低于训练得分,拟合问题求助

模型过拟合与泛化能力差问题的排查与解决

问题背景

拟合Logistic Regression模型时出现严重过拟合:网格搜索得到的最优参数在训练交叉验证中得分达到0.755,但测试集准确率仅0.49,混淆矩阵表现糟糕,且其他模型也存在同类问题。数据集共约900行,按30%比例拆分测试集,采用分层抽样。

原始代码

#importing libraries 
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.model_selection import GridSearchCV
from sklearn.preprocessing import OrdinalEncoder
from sklearn.preprocessing import LabelEncoder
from sklearn.preprocessing import MinMaxScaler

#creating test train splits
print(df['high_traffic'].value_counts())

y=df['high_traffic']
X= df.drop(["recipe","high_traffic"],axis=1)
X= pd.get_dummies(X,columns=['category'])

le = LabelEncoder()
X["servings"] = le.fit_transform(X["servings"])

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=3,stratify=y)


#scaling variables
scaler = StandardScaler()
scaled_train_X = scaler.fit_transform(X_train)


#PCA
pca=PCA()
pca.fit(scaled_train_X)
exp_variance = pca.explained_variance_ratio_
cum_exp_variance = np.cumsum(exp_variance)
print(cum_exp_variance)

pca = PCA(n_components=12,random_state=7)

train_pca = pca.fit_transform(scaled_train_X)
test_pca = pca.fit_transform(X_test)



    #logistic regression
    from sklearn.linear_model import LogisticRegression
    logreg = LogisticRegression(random_state=3)
    parameters = [    {'penalty' : ['l1', 'l2', 'elasticnet', 'none'],
    'C' : np.logspace(-4, 4, 20),
    'solver' : ['lbfgs','newton-cg','liblinear','sag','saga'],
    'max_iter' : [100, 1000,2500, 5000]
    }]
    grid_search = GridSearchCV(estimator = logreg, param_grid = parameters, scoring = 'accuracy', 
    cv = 5, verbose=1, refit=True)
    grid_search.fit(train_pca, y_train)

    #logreg.fit(train_pca,y_train)
    #pred_y_log = logreg.predict(test_pca)
    pred_y_log = grid_search.best_estimator_.predict(test_pca)
    print(grid_search.best_params_)
    print(grid_search.best_score_)

    from sklearn.metrics import classification_report
    class_log = classification_report(y_test,pred_y_log)
    print("Logistic Regression: \n", class_log)

    from sklearn.metrics import confusion_matrix
    print(confusion_matrix(y_test,pred_y_log))

原始运行结果

Fitting 5 folds for each of 1600 candidates, totalling 8000 fits
{'C': 0.23357214690901212, 'max_iter': 100, 'penalty': 'l1', 'solver': 'saga'}
0.7551049197672368
Logistic Regression:
accuracy 0.49
[[54 37]
[80 57]]


核心问题修复与优化方案

1. 修复数据预处理的致命错误

测试集未遵循"训练集拟合、测试集转换"的原则,导致分布偏移:

  • 错误代码:test_pca = pca.fit_transform(X_test)
  • 修正代码:
    # 用训练集拟合的scaler标准化测试集
    scaled_test_X = scaler.transform(X_test)
    # 用训练集拟合的PCA转换测试集
    test_pca = pca.transform(scaled_test_X)
    

2. 特征处理合理性调整

  • servings字段用LabelEncoder不合理:若该字段是数值型(如1人份、2人份),LabelEncoder会破坏数值的有序关系,建议改用MinMaxScaler或直接保留原始数值。
  • 检查类别分布:从混淆矩阵看测试集负类134条、正类94条,虽用了分层抽样,但仍需确认原始数据集是否存在更严重的类别不平衡。

3. 模型与评估逻辑优化

  • 替换评估指标:样本不平衡时准确率参考性差,网格搜索的scoring改为f1或roc_auc,模型评估改用F1-score、召回率/精确率。
  • 简化模型复杂度:当前网格搜索包含1600个参数组合,小数据集易过拟合,建议先固定参数组合(如l1对应saga/liblinear,l2对应lbfgs),减少候选数量。
  • 增强正则化:C值是正则化强度的倒数,当前最优C=0.23,可尝试更小的C值强化正则化,抑制过拟合。

4. 数据集与验证策略升级

  • 小数据集(900行)拆分30%测试集后仅270条样本,结果波动大,建议改用10折交叉验证评估泛化能力,或用ShuffleSplit多次拆分取平均结果。
  • 补充特征工程:尝试组合现有特征、生成统计特征,提升模型的有效学习维度。

内容的提问来源于stack exchange,提问作者Karolis Hematogenas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 23:00:03