如何用多评估器多特征组合自动筛选F1最优模型(不平衡数据集)
解决方案
核心思路
不用手动查看结果,我们可以把所有组合的性能指标(评估器名称、特征数量、F1分数、ROC AUC等)统一存储起来,然后按F1分数降序、再按ROC AUC降序排序,直接取出排名最靠前的组合即可。
修改后的完整代码
import numpy as np import pandas as pd from sklearn.feature_selection import RFE from sklearn.linear_model import LogisticRegression from sklearn.ensemble import RandomForestClassifier from sklearn.metrics import roc_auc_score, f1_score, roc_curve, precision_recall_curve, auc from sklearn.model_selection import cross_val_score import statsmodels.api as sm # 初始化评估器和特征数量列表 estimators = [('逻辑回归', LogisticRegression()), ('随机森林', RandomForestClassifier())] num_features_to_select = [4,5,7,9,11,15] # 存储所有组合的结果 results = [] for estimator_name, estimator in estimators: for n_features in num_features_to_select: # 创建RFE对象并拟合 rfe = RFE(estimator=estimator, n_features_to_select=n_features) rfe.fit(X_resampled, Y_resampled) # 获取选中的特征 selected_features = X_resampled.columns[rfe.support_] X_train_selected = X_resampled[selected_features] X_test_selected = X_test[selected_features] # 训练模型并预测(关闭迭代日志减少冗余输出) log_reg_model = sm.Logit(Y_resampled, X_train_selected).fit(disp=0) pred_test = log_reg_model.predict(X_test_selected) pred_test_1 = np.where(pred_test > 0.5, 1, 0) # 计算性能指标 logit_roc_auc = roc_auc_score(Y_test, pred_test) f1 = f1_score(Y_test, pred_test_1) precision, recall, _ = precision_recall_curve(Y_test, pred_test) prc_auc = auc(recall, precision) cv_scores = cross_val_score(estimator, X_resampled[selected_features], Y_resampled, cv=5) mean_cv_score = cv_scores.mean() # 将结果存入列表 results.append({ '评估器': estimator_name, '特征数量': n_features, 'F1分数': f1, 'ROC AUC': logit_roc_auc, 'PRC AUC': prc_auc, '平均CV分数': mean_cv_score, '选中特征': selected_features.tolist() }) # 转换为DataFrame方便排序和查看 results_df = pd.DataFrame(results) # 按F1分数降序,再按ROC AUC降序排序 sorted_results = results_df.sort_values(by=['F1分数', 'ROC AUC'], ascending=[False, False]) # 展示排名前5的组合(可按需调整数量) print("性能最优的组合排名:") print(sorted_results[['评估器', '特征数量', 'F1分数', 'ROC AUC']].head()) # 取出最优组合的详细信息 best_combination = sorted_results.iloc[0] print("\n最优组合详情:") print(f"评估器:{best_combination['评估器']}") print(f"特征数量:{best_combination['特征数量']}") print(f"测试集F1分数:{best_combination['F1分数']:.4f}") print(f"测试集ROC AUC:{best_combination['ROC AUC']:.4f}") print(f"选中的特征:{best_combination['选中特征']}")
关键改动说明
- 结果统一存储:用列表+字典的方式收集每个组合的所有关键信息,避免零散打印导致的信息混乱。
- 自动排序筛选:利用Pandas的排序功能,优先按F1分数从高到低排列,F1分数相同时再按ROC AUC排序,直接定位最优组合。
- 简化输出:只展示排名靠前的结果,无需手动翻阅所有输出内容。
- 细节优化:关闭
sm.Logit的迭代日志,减少不必要的冗余输出。
内容的提问来源于stack exchange,提问作者user20159866
相关产品推荐
相关产品推荐

