基于Bootstrap法的机器学习算法AUC置信区间求解报错求助
用Bootstrap计算机器学习算法AUC置信区间的索引错误解决方案
问题背景
使用含61个特征的个人医疗数据集(示例结构如下),尝试用Bootstrap法获取Logistic Regression等ML算法的AUC置信区间:
| Age | Female |
|---|---|
| 65 | 1 |
| 45 | 0 |
Logistic Regression的基础实现代码:
X = data_sevrage.drop(['Echec_sevrage'], axis=1) y = data_sevrage['Echec_sevrage'] X_train, X_test, y_train, y_test= train_test_split(X, y, test_size=0.25, random_state=0) lr = LogisticRegression(C=10 ,penalty='l1', solver= 'saga', max_iter=500).fit(X_train,y_train) score=roc_auc_score(y_test,lr.predict_proba(X_test)[:,1]) precision, recall, thresholds = precision_recall_curve(y_test, lr.predict_proba(X_test)[:,1]) auc_precision_recall = metrics.auc(recall, precision) y_pred = lr.predict(X_test) print('ROC AUC score :',score) print('auc_precision_recall :',auc_precision_recall)
尝试两种Bootstrap方法时均出现索引相关错误:
"None of [Int64Index([21, 22, 20, 31, 30, 13, 22, 1, 31, 3, 2, 9, 9, 18, 29, 30, 31,
31, 16, 11, 23, 7, 19, 10, 14, 5, 10, 25, 30, 24, 8, 20],
dtype='int64')] are in the [columns]"
'[6, 3, 12, 14, 10, 7, 9] not in index'
错误原因
这类索引错误的核心原因是:Bootstrap抽样时未同步处理特征矩阵X和标签y的行索引,或误将行位置索引当成列名/标签索引使用,导致样本与标签不匹配,或尝试访问不存在的列/索引。
正确的Bootstrap实现代码
使用sklearn.utils.resample工具可以自动保持X和y的行对应关系,避免手动处理索引的问题。以下是两种常用的Bootstrap实现方案:
方案1:基于固定训练/测试拆分的Bootstrap(推荐)
对训练集重复抽样训练模型,在固定测试集上评估,结果更稳定:
from sklearn.utils import resample import numpy as np from sklearn.metrics import roc_auc_score, precision_recall_curve, auc # Bootstrap参数设置 n_iterations = 1000 roc_auc_list = [] pr_auc_list = [] # 循环进行Bootstrap抽样 for idx in range(n_iterations): # 对训练集行进行有放回抽样,同步获取X和y的抽样样本 X_train_boot, y_train_boot = resample(X_train, y_train, random_state=idx) # 训练模型 lr_boot = LogisticRegression(C=10, penalty='l1', solver='saga', max_iter=500).fit(X_train_boot, y_train_boot) # 在固定测试集上计算ROC AUC y_proba = lr_boot.predict_proba(X_test)[:, 1] roc_auc = roc_auc_score(y_test, y_proba) roc_auc_list.append(roc_auc) # 计算PR AUC precision, recall, _ = precision_recall_curve(y_test, y_proba) pr_auc = auc(recall, precision) pr_auc_list.append(pr_auc) # 计算95%置信区间 roc_ci = np.percentile(roc_auc_list, [2.5, 97.5]) pr_ci = np.percentile(pr_auc_list, [2.5, 97.5]) print(f"ROC AUC 95%置信区间: [{roc_ci[0]:.4f}, {roc_ci[1]:.4f}]") print(f"PR AUC 95%置信区间: [{pr_ci[0]:.4f}, {pr_ci[1]:.4f}]")
方案2:对整个数据集的Bootstrap抽样
每次抽样后重新拆分训练/测试,模拟不同的数据拆分场景:
from sklearn.utils import resample import numpy as np from sklearn.metrics import roc_auc_score from sklearn.model_selection import train_test_split n_iterations = 1000 roc_auc_list = [] for idx in range(n_iterations): # 对整个数据集进行有放回抽样 data_boot = resample(data_sevrage, random_state=idx) # 拆分特征和标签 X_boot = data_boot.drop(['Echec_sevrage'], axis=1) y_boot = data_boot['Echec_sevrage'] # 拆分训练测试集 Xb_train, Xb_test, yb_train, yb_test = train_test_split(X_boot, y_boot, test_size=0.25, random_state=idx) # 训练并评估 lr_boot = LogisticRegression(C=10, penalty='l1', solver='saga', max_iter=500).fit(Xb_train, yb_train) roc_auc = roc_auc_score(yb_test, lr_boot.predict_proba(Xb_test)[:,1]) roc_auc_list.append(roc_auc) roc_ci = np.percentile(roc_auc_list, [2.5, 97.5]) print(f"ROC AUC 95%置信区间: [{roc_ci[0]:.4f}, {roc_ci[1]:.4f}]")
关键注意事项
- 必须同步抽样X和y的行,确保每个样本的特征与标签一一对应
- 优先使用
sklearn.utils.resample工具,避免手动处理索引时出现的匹配错误 - 如果手动实现抽样,需使用行的位置索引(
iloc),而非标签索引,例如:# 手动抽样示例:生成行位置索引 sample_idx = np.random.choice(len(X_train), len(X_train), replace=True) X_train_boot = X_train.iloc[sample_idx] y_train_boot = y_train.iloc[sample_idx]
内容的提问来源于stack exchange,提问作者Romain LOMBARDI
相关产品推荐
相关产品推荐

