You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Bootstrap法的机器学习算法AUC置信区间求解报错求助

用Bootstrap计算机器学习算法AUC置信区间的索引错误解决方案

问题背景

使用含61个特征的个人医疗数据集(示例结构如下),尝试用Bootstrap法获取Logistic Regression等ML算法的AUC置信区间:

AgeFemale
651
450

Logistic Regression的基础实现代码:

X = data_sevrage.drop(['Echec_sevrage'], axis=1)
y = data_sevrage['Echec_sevrage']

X_train, X_test, y_train, y_test= train_test_split(X, y, test_size=0.25, random_state=0)

lr = LogisticRegression(C=10 ,penalty='l1', solver= 'saga', max_iter=500).fit(X_train,y_train)
score=roc_auc_score(y_test,lr.predict_proba(X_test)[:,1])
precision, recall, thresholds = precision_recall_curve(y_test, lr.predict_proba(X_test)[:,1])
auc_precision_recall = metrics.auc(recall, precision)
y_pred = lr.predict(X_test)
print('ROC AUC score :',score)
print('auc_precision_recall :',auc_precision_recall)

尝试两种Bootstrap方法时均出现索引相关错误:

"None of [Int64Index([21, 22, 20, 31, 30, 13, 22, 1, 31, 3, 2, 9, 9, 18, 29, 30, 31,
31, 16, 11, 23, 7, 19, 10, 14, 5, 10, 25, 30, 24, 8, 20],
dtype='int64')] are in the [columns]"

'[6, 3, 12, 14, 10, 7, 9] not in index'


错误原因

这类索引错误的核心原因是:Bootstrap抽样时未同步处理特征矩阵X和标签y的行索引,或误将行位置索引当成列名/标签索引使用,导致样本与标签不匹配,或尝试访问不存在的列/索引。


正确的Bootstrap实现代码

使用sklearn.utils.resample工具可以自动保持X和y的行对应关系,避免手动处理索引的问题。以下是两种常用的Bootstrap实现方案:

方案1:基于固定训练/测试拆分的Bootstrap(推荐)

对训练集重复抽样训练模型,在固定测试集上评估,结果更稳定:

from sklearn.utils import resample
import numpy as np
from sklearn.metrics import roc_auc_score, precision_recall_curve, auc

# Bootstrap参数设置
n_iterations = 1000
roc_auc_list = []
pr_auc_list = []

# 循环进行Bootstrap抽样
for idx in range(n_iterations):
    # 对训练集行进行有放回抽样,同步获取X和y的抽样样本
    X_train_boot, y_train_boot = resample(X_train, y_train, random_state=idx)
    # 训练模型
    lr_boot = LogisticRegression(C=10, penalty='l1', solver='saga', max_iter=500).fit(X_train_boot, y_train_boot)
    # 在固定测试集上计算ROC AUC
    y_proba = lr_boot.predict_proba(X_test)[:, 1]
    roc_auc = roc_auc_score(y_test, y_proba)
    roc_auc_list.append(roc_auc)
    # 计算PR AUC
    precision, recall, _ = precision_recall_curve(y_test, y_proba)
    pr_auc = auc(recall, precision)
    pr_auc_list.append(pr_auc)

# 计算95%置信区间
roc_ci = np.percentile(roc_auc_list, [2.5, 97.5])
pr_ci = np.percentile(pr_auc_list, [2.5, 97.5])

print(f"ROC AUC 95%置信区间: [{roc_ci[0]:.4f}, {roc_ci[1]:.4f}]")
print(f"PR AUC 95%置信区间: [{pr_ci[0]:.4f}, {pr_ci[1]:.4f}]")

方案2:对整个数据集的Bootstrap抽样

每次抽样后重新拆分训练/测试,模拟不同的数据拆分场景:

from sklearn.utils import resample
import numpy as np
from sklearn.metrics import roc_auc_score
from sklearn.model_selection import train_test_split

n_iterations = 1000
roc_auc_list = []

for idx in range(n_iterations):
    # 对整个数据集进行有放回抽样
    data_boot = resample(data_sevrage, random_state=idx)
    # 拆分特征和标签
    X_boot = data_boot.drop(['Echec_sevrage'], axis=1)
    y_boot = data_boot['Echec_sevrage']
    # 拆分训练测试集
    Xb_train, Xb_test, yb_train, yb_test = train_test_split(X_boot, y_boot, test_size=0.25, random_state=idx)
    # 训练并评估
    lr_boot = LogisticRegression(C=10, penalty='l1', solver='saga', max_iter=500).fit(Xb_train, yb_train)
    roc_auc = roc_auc_score(yb_test, lr_boot.predict_proba(Xb_test)[:,1])
    roc_auc_list.append(roc_auc)

roc_ci = np.percentile(roc_auc_list, [2.5, 97.5])
print(f"ROC AUC 95%置信区间: [{roc_ci[0]:.4f}, {roc_ci[1]:.4f}]")

关键注意事项

  • 必须同步抽样X和y的行,确保每个样本的特征与标签一一对应
  • 优先使用sklearn.utils.resample工具,避免手动处理索引时出现的匹配错误
  • 如果手动实现抽样,需使用行的位置索引(iloc),而非标签索引,例如:
    # 手动抽样示例:生成行位置索引
    sample_idx = np.random.choice(len(X_train), len(X_train), replace=True)
    X_train_boot = X_train.iloc[sample_idx]
    y_train_boot = y_train.iloc[sample_idx]
    

内容的提问来源于stack exchange,提问作者Romain LOMBARDI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 05:25:24