You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Sklearn中BaggingClassifier结合预处理Pipeline使用报错求助

问题根源

错误的核心原因是BaggingClassifier在bootstrap采样后,会将输入的pandas DataFrame转换为numpy数组,而你的预处理流程(ColumnTransformer)是通过字符串列名来指定特征列的——这种方式仅支持pandas DataFrame输入,当输入变成numpy数组时,数组没有columns属性,就触发了报错。

单独运行Pipeline时,输入始终是DataFrame,所以列名选择能正常工作;但Bagging内部采样后传递给基学习器的是numpy数组,导致ColumnTransformer无法识别列名。

两种解决方案

方案1:改用列索引指定特征(推荐)

将原来用字符串列名定义特征的方式,改成对应列的索引位置,这样不管输入是DataFrame还是numpy数组,ColumnTransformer都能正常定位特征列。

修改代码中特征定义的部分:

# 用列索引替代列名
numerical_features = [X.columns.get_loc(col) for col in ['age', 'fare']] 
categorical_features = [X.columns.get_loc(col) for col in ['sex', 'deck', 'alone']] 
other_features = [X.columns.get_loc(col) for col in ['pclass']]

其余代码保持不变,重新运行即可。

方案2:强制Bagging传递DataFrame给基学习器

在Pipeline的最前端添加一个FunctionTransformer,将Bagging传递的numpy数组转回pandas DataFrame,确保后续ColumnTransformer能识别列名。

修改预处理流程部分:

from sklearn.preprocessing import FunctionTransformer

# 定义转换函数,将数组转回DataFrame
def array_to_dataframe(X):
    return pd.DataFrame(X, columns=X_train.columns)

# 在预处理前添加转DataFrame的步骤
preprocessor = make_pipeline(
    FunctionTransformer(array_to_dataframe),
    make_column_transformer(
        (numerical_pipeline, numerical_features),
        (categorical_pipeline, categorical_features),
        (other_pipeline, other_features),
    )
)

这种方法需要确保训练和测试时的列顺序一致,适合必须保留列名逻辑的场景。

验证修改

修改后重新运行model_bagging.fit(X_train,y_train)和score方法,就能正常执行装袋训练和评估了。

内容的提问来源于stack exchange,提问作者Guigui

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 18:14:49