You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GridSearchCV评分参数疑问:交叉验证与测试集Precision差异解析

随机森林GridSearchCV交叉验证精度与测试集精度差异原因分析

问题描述

我在用GridSearchCV对随机森林分类器调参时,将scoring参数设为precision,得到的最佳交叉验证精度为0.9639,但用该最佳参数训练模型后,在测试集上的精度表现远低于预期,请问这种差异的原因是什么?

相关代码

from sklearn.model_selection import GridSearchCV

params_rf = { 'criterion': ['gini','entropy'],           
     'min_samples_leaf': [1,3,5,10,20,25,30,35,40,45,50]
                 }
rf = RandomForestClassifier(n_estimators=2500,
                      random_state=SEED,
                      max_features='sqrt',
                      bootstrap=True,
                      oob_score=True,
                      n_jobs=-1,
                      class_weight={1:3},
                      warm_start=True,
                      refit=True,
                      return_train_score=True)

prec = make_scorer(precision_score)
                    
grid_rf = GridSearchCV(estimator = rf, 
param_grid=params_rf,scoring=prec,cv=10,n_jobs=-1,
verbose=True)

grid_rf.fit(X_resampled,y_resampled)

y_pred = grid_rf.predict(X_test)

best_hyperparams = grid_rf.best_params_
best_score = grid_rf.best_score_
best_estimator = grid_rf.best_estimator_
print('Best hyperparameters:\n', best_hyperparams)
print('Best score:\n', best_score.round(4))
print('Best estimator:\n', best_estimator)


# 输出结果
# Fitting 10 folds for each of 22 candidates, totalling 220 fits
# Best hyperparameters:
#  {'criterion': 'entropy', 'min_samples_leaf': 1}
# Best score:
#  0.9639
# Best estimator:
#  RandomForestClassifier(class_weight={1: 3}, 
# criterion='entropy',
#                n_estimators=2500, n_jobs=-1, oob_score=True,
#                random_state=121864, warm_start=True)


# 使用最佳参数训练模型(A)
rf_A = RandomForestClassifier(class_weight={1: 3}, 
criterion='entropy',
               n_estimators=2500, n_jobs=-1, oob_score=True,
               random_state=121864, 
warm_start=True,min_samples_leaf=1                          
                      )
          
rf_A.fit(X_resampled,y_resampled)
y_pred_A=(rf_A.predict(X_test))
importances_rf_A = pd.Series(rf_A.feature_importances_, index 
= Features.columns)

# 混淆矩阵热力图
matrix =confusion_matrix(y_test,y_pred_A).round(2)
text = np.array([['True Positive', 'False Negative'],
        ['False Positive', 'True Negative']])

formatted_text = (np.asarray(["{0}\n{1:.0f}".format(
text, matrix) for text, matrix in zip(text.flatten(), 
matrix.flatten())])).reshape(2,2)

fig, ax = plt.subplots(figsize=(7,2))
sns.set(font_scale=1.3)
ax = sns.heatmap(matrix, annot=formatted_text, fmt="", 
linewidth=1,cbar=False)
ax.set_title('Confusion Matrix', size = 18)
plt.show()

target_names = ['Fully Paid', 'Not Fully Paid']
print(classification_report(y_test,y_pred_A,
target_names=target_names,zero_division=1))

# 特征重要性条形图
sorted_importances_rf_A = importances_rf_A.sort_values()
sorted_importances_descend_A = importances_rf_A.sort_values(ascending=False)
plt.figure(figsize=(20, 15))
sorted_importances_rf_A.plot(kind='barh',grid=True)
plt.show()
print(sorted_importances_descend_A.cumsum())

测试集分类报告

precision    recall  f1-score   support

Fully Paid       0.87      0.65      0.75      1609
Not Fully Paid       0.22      0.51      0.31       307

    accuracy                           0.63      1916
   macro avg       0.55      0.58      0.53      1916
weighted avg       0.77      0.63      0.68      1916

差异原因分析

  • 数据分布不一致:GridSearchCV的交叉验证基于重采样后的X_resampled,y_resampled,这类数据为解决类别不平衡问题(如过采样少数类),样本分布更均衡;而测试集X_test,y_test是真实的不平衡分布(Fully Paid样本数是Not Fully Paid的5倍多)。模型在均衡数据上训练时,少数类的预测精度被高估,但在真实不平衡场景下,大量负类样本会产生更多假阳性,拉低整体精度。
  • 精度指标的类别指向性:make_scorer(precision_score)默认在二分类任务中计算**正类(类别1,即Not Fully Paid)**的精度。交叉验证时重采样后正类样本充足,模型能精准识别正类,得到0.9639的高精度;但测试集中正类样本少,模型容易将负类误判为正类,导致正类精度骤降至0.22,若误将交叉验证的正类精度和测试集的加权平均精度/负类精度对比,会产生认知偏差。
  • 模型过拟合风险:最佳参数中min_samples_leaf=1,允许决策树在单个样本上分裂,会让模型过度拟合重采样数据中的噪声(包括合成样本的人工特征)。训练数据上表现极佳,但面对分布不同的真实测试集时,泛化能力大幅下降,精度自然降低。
  • 交叉验证分数高估:重采样生成的合成样本和原始样本高度相似,破坏了交叉验证中样本的独立性,导致交叉验证分数无法真实反映模型泛化能力,出现“分数虚高”的情况。

内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 18:15:47