Python sklearn构建pipeline时提示pipeline未定义的解决方法
报错修复方案:
name 'pipeline' is not defined 报错原因
触发这个错误有两个直接问题:
- 代码只从
sklearn.pipeline导入了FeatureUnion,没有导入流水线类Pipeline - scikit-learn的流水线类采用大驼峰命名,名称为大写开头的
Pipeline,代码里写成了全小写的pipeline,Python大小写敏感,无法识别这个未定义的名称。
修复步骤
- 补全导入语句,在代码开头的导入部分加入
Pipeline的引入,修改第一行导入为:
from sklearn.pipeline import FeatureUnion, Pipeline
- 将代码中两处创建数值、分类流水线时写的小写
pipeline(,全部替换为大写开头的Pipeline(。
如果你使用的是0.22及以上版本的scikit-learn,还需要处理两个版本兼容问题,避免后续出现其他报错:
- 旧版位于
sklearn.preprocessing的Imputer已经被迁移到sklearn.impute模块,重命名为SimpleImputer,需要额外导入from sklearn.impute import SimpleImputer,并将代码中Imputer(strategy="median")替换为SimpleImputer(strategy="median")- 新版
LabelBinarizer默认返回稀疏矩阵,会导致特征拼接失败,可以在初始化时加入参数sparse_output=False(0.24之前的旧版本用参数sparse=False)
修正后的完整参考代码
# 补全所有需要的导入 from sklearn.pipeline import FeatureUnion, Pipeline from sklearn.impute import SimpleImputer from sklearn.preprocessing import StandardScaler, LabelBinarizer # 注:DataFrameSelector、CombinedAttributesAdder是书中自定义的转换器,需要提前按照书中示例定义好才能运行 num_attribs = list(housing_num) cat_attribs = ["ocean_proximity"] num_pipeline = Pipeline([ ('selector', DataFrameSelector(num_attribs)), ('imputer', SimpleImputer(strategy="median")), ('attribs_adder', CombinedAttributesAdder()), ('std_scaler', StandardScaler()), ]) cat_pipeline = Pipeline([ ('selector', DataFrameSelector(cat_attribs)), ('label_binarizer', LabelBinarizer(sparse_output=False)), ]) full_pipeline = FeatureUnion(transformer_list=[ ("num_pipeline", num_pipeline), ("cat_pipeline", cat_pipeline), ])
内容的提问来源于stack exchange,提问作者Nishant Chaudhary
相关产品推荐
相关产品推荐

