You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否用scipy.stats.bootstrap计算回归/分类任务中特征权重的置信区间?

使用scipy.stats.bootstrap计算特征权重的置信区间

是的,scipy.stats.bootstrap完全可以实现用bootstrap方法计算分类/回归模型特征权重置信区间的需求。下面结合你提供的LDA分类任务示例,给出具体实现步骤:

核心思路

对原始数据集(X和y)进行多次有放回重采样,每次采样后拟合模型并提取特征权重,最终基于所有重采样得到的权重分布计算置信区间。

完整代码实现

import numpy as np
from sklearn.discriminant_analysis import LinearDiscriminantAnalysis
from scipy.stats import bootstrap

# 原始数据集
X = np.array([[-1, -1], [-2, -1], [-3, -2], [1, 1], [2, 1], [3, 2]])
y = np.array([1, 1, 1, 2, 2, 2])

# 定义bootstrap所需的统计量函数:输入采样后的X、y,返回特征权重
def get_feature_weights(sample_X, sample_y):
    clf = LinearDiscriminantAnalysis()
    clf.fit(sample_X, sample_y)
    # 返回扁平化的系数数组,适配scipy.bootstrap的输入要求
    return clf.coef_.flatten()

# 将X和y打包成bootstrap需要的样本组
data = (X, y)

# 执行bootstrap计算
bootstrap_result = bootstrap(
    data,
    statistic=get_feature_weights,
    n_resamples=1000,
    confidence_level=0.95,
    method='percentile'  # 可选'basic'/'bca'等其他置信区间计算方法
)

# 输出特征权重的95%置信区间
print("特征权重的95%置信区间:")
for i, (lower, upper) in enumerate(zip(bootstrap_result.confidence_interval.low, bootstrap_result.confidence_interval.high)):
    print(f"特征{i+1}: [{lower:.4f}, {upper:.4f}]")

# 如需查看所有重采样的权重分布,可取消注释
# print(bootstrap_result.bootstrap_distribution)

关键说明

  • statistic参数需传入一个与data参数匹配的函数,这里的get_feature_weights接收采样后的X和y,返回模型的特征权重。
  • method参数的可选值中,'bca'(偏差校正加速法)的置信区间准确性更高,但计算成本略高于'percentile'和'basic'。
  • 回归任务的实现逻辑完全一致,只需替换模型(如线性回归、随机森林回归)并提取对应的特征权重(如coef_或特征重要性)即可。

内容的提问来源于stack exchange,提问作者Johannes Wiesner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 20:04:53