使用sklearn make_pipeline处理泰坦尼克数据集遇TypeError求助
泰坦尼克数据集预处理管道构建错误排查
问题描述
基于泰坦尼克数据集使用sklearn的make_pipeline构建预处理管道时,调用preprocessing.fit_transform(titanic_data)触发TypeError,提示管道中的估计器不符合要求。
实现代码
def sum_relatives(X): X_copy = X.copy() X_copy['total_relatives'] = X_copy['SibSp'] + X_copy['Parch'] return X_copy class_order = [[1, 2, 3]] ord_pipeline = make_pipeline( OrdinalEncoder(categories=class_order) ) def age_transformer(X): X_copy = X.copy() for index, row in self.median_age_by_class.iterrows(): class_value = row['Pclass'] median_age = row['median_age'] X_copy.loc[X_copy['Pclass'] == class_value, 'Age'] = X_copy.loc[X_copy['Pclass'] == class_value, 'Age'].fillna(median_age) bins = [0, 10, 20, 30, 40, 50, 60, 70, 100] X_copy['age_interval'] = pd.cut(X_copy['Age'], bins=bins) return X_copy def age_processor(): return make_pipeline( FunctionTransformer(age_transformer), ) total_relatives_pipeline = make_pipeline( FunctionTransformer(sum_relatives) ) cat_pipeline = make_pipeline( OneHotEncoder(handle_unknown="ignore") ) num_pipeline = make_pipeline([ StandardScaler() ]) preprocessing = ColumnTransformer([ ("ord", ord_pipeline, ['Pclass']), ("age_processing", age_processor(), ['Pclass', 'Age']), ("total_relatives", total_relatives_pipeline, ['SibSp', 'Parch']), ("cat", cat_pipeline, ['Sex', 'Embarked', 'traveling_category', 'age_interval']), ("num", num_pipeline, ['Fare']), ])
错误信息
TypeError: All estimators should implement fit and transform, or can be 'drop' or 'passthrough' specifiers. 'Pipeline(steps=[('list', [('scaler', StandardScaler())])])' (type <class 'sklearn.pipeline.Pipeline'>) doesn't.
错误分析与修复步骤
1. 直接触发错误的原因:num_pipeline参数格式错误
make_pipeline要求直接传入转换器实例,不需要用列表包裹。你写成了make_pipeline([StandardScaler()]),导致Pipeline将这个列表作为一个名为list的步骤——而列表不是符合sklearn规范的估计器(没有fit/transform方法),这就是报错的核心原因。
修复代码:
num_pipeline = make_pipeline( StandardScaler() )
2. 潜在运行错误:age_transformer中的self引用无效
age_transformer是普通函数,不是类方法,代码里的self.median_age_by_class没有定义,运行时会触发NameError,必须修正:
修复方案一:通过参数传递中位数数据
def age_transformer(X, median_age_by_class): X_copy = X.copy() for index, row in median_age_by_class.iterrows(): class_value = row['Pclass'] median_age = row['median_age'] X_copy.loc[X_copy['Pclass'] == class_value, 'Age'] = X_copy.loc[X_copy['Pclass'] == class_value, 'Age'].fillna(median_age) bins = [0, 10, 20, 30, 40, 50, 60, 70, 100] X_copy['age_interval'] = pd.cut(X_copy['Age'], bins=bins) return X_copy # 先计算各舱位的年龄中位数 median_age_by_class = titanic_data.groupby('Pclass')['Age'].median().reset_index(name='median_age') def age_processor(): return make_pipeline( FunctionTransformer(age_transformer, kw_args={'median_age_by_class': median_age_by_class}) )
修复方案二:自定义符合sklearn规范的转换器类
from sklearn.base import BaseEstimator, TransformerMixin class AgeTransformer(BaseEstimator, TransformerMixin): def fit(self, X, y=None): # 在fit阶段计算并保存各舱位年龄中位数 self.median_age_by_class = X.groupby('Pclass')['Age'].median().reset_index(name='median_age') return self def transform(self, X): X_copy = X.copy() for index, row in self.median_age_by_class.iterrows(): class_value = row['Pclass'] median_age = row['median_age'] X_copy.loc[X_copy['Pclass'] == class_value, 'Age'] = X_copy.loc[X_copy['Pclass'] == class_value, 'Age'].fillna(median_age) bins = [0, 10, 20, 30, 40, 50, 60, 70, 100] X_copy['age_interval'] = pd.cut(X_copy['Age'], bins=bins) return X_copy def age_processor(): return make_pipeline( AgeTransformer() )
3. 额外注意:ColumnTransformer的列匹配问题
cat步骤中指定的age_interval是age_processing步骤生成的新列,原始数据中不存在,默认情况下ColumnTransformer不会自动传递跨步骤的列。可以通过设置remainder='passthrough'保留处理后的所有列,再调整步骤顺序或单独处理age_interval:
preprocessing = ColumnTransformer([ ("ord", ord_pipeline, ['Pclass']), ("total_relatives", total_relatives_pipeline, ['SibSp', 'Parch']), ("age_processing", age_processor(), ['Pclass', 'Age']), ("cat", cat_pipeline, ['Sex', 'Embarked', 'traveling_category']), ("num", num_pipeline, ['Fare']), ], remainder='passthrough')
内容的提问来源于stack exchange,提问作者Silvio sjsj
相关产品推荐
相关产品推荐

