ColumnTransformer缺失_name_to_fitted_passthrough属性,Flask预测报错
问题:Flask调用sklearn ColumnTransformer.transform()报属性错误
问题详情
基于Python 3.8虚拟环境开发MLOps项目,使用最新版scikit-learn构建ColumnTransformer预处理对象(包含数值、分类两条管道),训练后保存为preprocessor.pkl。在Flask应用中调用preprocessor.transform()做预测时,抛出错误:
'ColumnTransformer' object has no attribute '_name_to_fitted_passthrough'
相关代码
from dataclasses import dataclass import os import pandas as pd import numpy as np from sklearn.pipeline import Pipeline from sklearn.impute import SimpleImputer from sklearn.preprocessing import StandardScaler, OneHotEncoder from sklearn.compose import ColumnTransformer import logging import sys # 项目自定义异常类与工具函数 class CustomException(Exception): def __init__(self, e, sys=None): super().__init__(str(e)) self.sys = sys def save_objects(file_path, obj): import pickle with open(file_path, 'wb') as f: pickle.dump(obj, f) @dataclass class DataTransformationConfig: preprocessor_obj_file_path=os.path.join('artifacts','preprocessor.pkl') class DataTransformation: def __init__(self): self.data_tranformation_config=DataTransformationConfig() def get_transfromation_object(self): try: numerical_column = ["writing_score", "reading_score"] categorical_columns = [ "gender", "race_ethnicity", "parental_level_of_education", "lunch", "test_preparation_course", ] num_pipeline = Pipeline( steps=[ ('imputer',SimpleImputer(strategy='median')), ('scaler',StandardScaler()) ] ) cat_pipeline = Pipeline( steps=[ ('imputer',SimpleImputer(strategy='most_frequent')), ('one_hot_encoder',OneHotEncoder()), ('scaler',StandardScaler(with_mean=False)) ] ) logging.info(f'Categorical columns:{categorical_columns}') logging.info(f'Numeric columns:{numerical_column}') preprocessor=ColumnTransformer( [ ('num_pipeline',num_pipeline,numerical_column), ('cat_pipeline',cat_pipeline,categorical_columns) ] ) logging.info(f"Whole pipeline{preprocessor}") return preprocessor except Exception as e: raise CustomException(e) def initiate_data_tranformation(self,train_path,test_path): try: train_df=pd.read_csv(train_path) test_df=pd.read_csv(test_path) logging.info('Read train and test data completed') logging.info('obtaining preprocessing object') preprocessor_obj=self.get_transfromation_object() target_column_name="math_score" input_feature_train_df=train_df.drop(columns=[target_column_name],axis=1) target_feature_train_df=train_df[target_column_name] input_feature_test_df=test_df.drop(columns=[target_column_name],axis=1) target_feature_test_df=test_df[target_column_name] logging.info(f'Applying preprocessing object on train and test dataframe.') input_feature_train_arr = preprocessor_obj.fit_transform(input_feature_train_df) input_feature_test_arr = preprocessor_obj.transform(input_feature_test_df) logging.info(f"train shape{input_feature_train_arr.shape} and test shape {input_feature_test_arr.shape}") train_arr=np.c_[input_feature_train_arr,np.array(target_feature_train_df)] test_arr=np.c_[input_feature_test_arr,np.array(target_feature_test_df)] logging.info(f'saved preprocessing object') save_objects( file_path=self.data_tranformation_config.preprocessor_obj_file_path, obj=preprocessor_obj ) return (train_arr, test_arr, self.data_tranformation_config.preprocessor_obj_file_path) except Exception as e: raise CustomException(e,sys)
解决方案
这个错误的核心是训练环境与Flask运行环境的scikit-learn版本不匹配:
_name_to_fitted_passthrough是scikit-learn 1.2及以上版本为ColumnTransformer新增的属性。如果训练时用了高版本sklearn,而Flask环境用了更低版本,加载序列化后的模型后调用transform()就会因低版本类缺少该属性报错。
解决步骤:
- 统一依赖版本:在训练环境执行
pip freeze > requirements.txt,导出所有依赖的精确版本;在Flask部署环境执行pip install -r requirements.txt,确保两者依赖完全一致。 - 重新生成预处理对象:版本统一后,重新运行数据转换流程,生成新的
preprocessor.pkl替换旧文件。 - 验证功能:在Flask环境加载新的预处理对象,测试
transform()方法是否正常执行。
另外,检查save_objects函数的序列化逻辑,确保使用标准pickle工具,无自定义的不兼容序列化逻辑。
内容的提问来源于stack exchange,提问作者VIJAY KUMAR
相关产品推荐
相关产品推荐

