如何为Python Pipeline添加交叉验证并报告模型评分?
为机器学习Pipeline添加交叉验证的实现方案
你当前的Pipeline代码如下:
pipe = Pipeline(steps=[ ("fasttext", FastTextVectorTransformer()), ("umap", umap_skl(n_components=8)), ("classifier", HistGradientBoostingClassifier()) ]) pipe.fit(X_train, y_train) print(f'Fbeta_score:{fbeta_score(y[test_index], pipe.predict(X[test_index]), average=None, beta=1.4)}') print(f'Accuracy:{accuracy_score(y[test_index], pipe.predict(X[test_index]))}') print(f'Precission:{precision_score(y_test, pipe.predict(X[test_index]))}') print(f'Recall:{recall_score(y_test, pipe.predict(X_test))}') print(f'Roc_AUC:{roc_auc_score(y_test, pipe.predict(X_test))}')
要改用交叉验证评估模型性能,可借助scikit-learn的cross_val_score或cross_validate工具,以下是具体实现:
步骤1:导入所需模块
from sklearn.model_selection import cross_val_score, cross_validate from sklearn.metrics import make_scorer, fbeta_score, accuracy_score, precision_score, recall_score, roc_auc_score
步骤2:自定义F-beta评分器
由于cross_validate没有内置beta=1.4的F-beta评分,需要先自定义评分器:
# 自定义beta=1.4的F-beta评分器 fbeta_scorer = make_scorer(fbeta_score, beta=1.4, average=None) # 若需要宏平均/微平均,可调整average参数,比如average='macro'
步骤3:使用cross_validate执行多指标交叉验证
cross_validate支持同时计算多个评估指标,更匹配你的需求。直接用训练集数据(X_train, y_train)执行交叉验证,避免数据泄露:
# 定义需要评估的指标集合 scoring = { 'fbeta': fbeta_scorer, 'accuracy': make_scorer(accuracy_score), 'precision': make_scorer(precision_score), 'recall': make_scorer(recall_score), 'roc_auc': make_scorer(roc_auc_score) } # 执行5折交叉验证(可通过cv参数调整折数) cv_results = cross_validate(pipe, X_train, y_train, cv=5, scoring=scoring) # 输出各指标的交叉验证结果 print("交叉验证各指标得分:") print(f"F-beta得分(beta=1.4):{cv_results['test_fbeta']}") print(f"F-beta均值:{cv_results['test_fbeta'].mean(axis=0)},标准差:{cv_results['test_fbeta'].std(axis=0)}") print(f"Accuracy得分:{cv_results['test_accuracy']}") print(f"Accuracy均值:{cv_results['test_accuracy'].mean()},标准差:{cv_results['test_accuracy'].std()}") print(f"Precision得分:{cv_results['test_precision']}") print(f"Precision均值:{cv_results['test_precision'].mean()},标准差:{cv_results['test_precision'].std()}") print(f"Recall得分:{cv_results['test_recall']}") print(f"Recall均值:{cv_results['test_recall'].mean()},标准差:{cv_results['test_recall'].std()}") print(f"ROC-AUC得分:{cv_results['test_roc_auc']}") print(f"ROC-AUC均值:{cv_results['test_roc_auc'].mean()},标准差:{cv_results['test_roc_auc'].std()}")
补充说明
- 交叉验证无需提前拆分测试集,
cross_validate会自动将训练集划分为训练子集和验证子集,每折都完成"训练-评估"流程 - 若只需要单一指标评估,可使用
cross_val_score简化操作,例如:
accuracy_scores = cross_val_score(pipe, X_train, y_train, cv=5, scoring='accuracy') print(f"Accuracy交叉验证得分:{accuracy_scores},均值:{accuracy_scores.mean()}")
内容的提问来源于stack exchange,提问作者Maite89
相关产品推荐
相关产品推荐

