You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用sklearn make_pipeline处理泰坦尼克数据集遇TypeError求助

泰坦尼克数据集预处理管道构建错误排查

问题描述

基于泰坦尼克数据集使用sklearn的make_pipeline构建预处理管道时,调用preprocessing.fit_transform(titanic_data)触发TypeError,提示管道中的估计器不符合要求。

实现代码

def sum_relatives(X):
    X_copy = X.copy()
    X_copy['total_relatives'] = X_copy['SibSp'] + X_copy['Parch']
    return X_copy

class_order = [[1, 2, 3]]

ord_pipeline = make_pipeline(
    OrdinalEncoder(categories=class_order)    
    )

def age_transformer(X):
    X_copy = X.copy()
    for index, row in self.median_age_by_class.iterrows():
        class_value = row['Pclass']
        median_age = row['median_age']
        X_copy.loc[X_copy['Pclass'] == class_value, 'Age'] = X_copy.loc[X_copy['Pclass'] == class_value, 'Age'].fillna(median_age)
    bins = [0, 10, 20, 30, 40, 50, 60, 70, 100]
    X_copy['age_interval'] = pd.cut(X_copy['Age'], bins=bins)
    return X_copy

def age_processor():
    return make_pipeline(
        FunctionTransformer(age_transformer),
)

total_relatives_pipeline = make_pipeline(
    FunctionTransformer(sum_relatives)
)

cat_pipeline = make_pipeline(
    OneHotEncoder(handle_unknown="ignore")
)

num_pipeline = make_pipeline([
        StandardScaler()
])

preprocessing = ColumnTransformer([
    ("ord", ord_pipeline, ['Pclass']),
    ("age_processing", age_processor(), ['Pclass', 'Age']),
    ("total_relatives", total_relatives_pipeline, ['SibSp', 'Parch']),
    ("cat", cat_pipeline, ['Sex', 'Embarked', 'traveling_category', 'age_interval']),
    ("num", num_pipeline, ['Fare']),
])

错误信息

TypeError: All estimators should implement fit and transform, or can be 'drop' or 'passthrough' specifiers. 'Pipeline(steps=[('list', [('scaler', StandardScaler())])])' (type <class 'sklearn.pipeline.Pipeline'>) doesn't.

错误分析与修复步骤

1. 直接触发错误的原因:num_pipeline参数格式错误

make_pipeline要求直接传入转换器实例,不需要用列表包裹。你写成了make_pipeline([StandardScaler()]),导致Pipeline将这个列表作为一个名为list的步骤——而列表不是符合sklearn规范的估计器(没有fit/transform方法),这就是报错的核心原因。

修复代码:

num_pipeline = make_pipeline(
    StandardScaler()
)

2. 潜在运行错误:age_transformer中的self引用无效

age_transformer是普通函数,不是类方法,代码里的self.median_age_by_class没有定义,运行时会触发NameError,必须修正:

修复方案一:通过参数传递中位数数据

def age_transformer(X, median_age_by_class):
    X_copy = X.copy()
    for index, row in median_age_by_class.iterrows():
        class_value = row['Pclass']
        median_age = row['median_age']
        X_copy.loc[X_copy['Pclass'] == class_value, 'Age'] = X_copy.loc[X_copy['Pclass'] == class_value, 'Age'].fillna(median_age)
    bins = [0, 10, 20, 30, 40, 50, 60, 70, 100]
    X_copy['age_interval'] = pd.cut(X_copy['Age'], bins=bins)
    return X_copy

# 先计算各舱位的年龄中位数
median_age_by_class = titanic_data.groupby('Pclass')['Age'].median().reset_index(name='median_age')

def age_processor():
    return make_pipeline(
        FunctionTransformer(age_transformer, kw_args={'median_age_by_class': median_age_by_class})
    )

修复方案二:自定义符合sklearn规范的转换器类

from sklearn.base import BaseEstimator, TransformerMixin

class AgeTransformer(BaseEstimator, TransformerMixin):
    def fit(self, X, y=None):
        # 在fit阶段计算并保存各舱位年龄中位数
        self.median_age_by_class = X.groupby('Pclass')['Age'].median().reset_index(name='median_age')
        return self
    
    def transform(self, X):
        X_copy = X.copy()
        for index, row in self.median_age_by_class.iterrows():
            class_value = row['Pclass']
            median_age = row['median_age']
            X_copy.loc[X_copy['Pclass'] == class_value, 'Age'] = X_copy.loc[X_copy['Pclass'] == class_value, 'Age'].fillna(median_age)
        bins = [0, 10, 20, 30, 40, 50, 60, 70, 100]
        X_copy['age_interval'] = pd.cut(X_copy['Age'], bins=bins)
        return X_copy

def age_processor():
    return make_pipeline(
        AgeTransformer()
    )

3. 额外注意:ColumnTransformer的列匹配问题

cat步骤中指定的age_interval是age_processing步骤生成的新列,原始数据中不存在,默认情况下ColumnTransformer不会自动传递跨步骤的列。可以通过设置remainder='passthrough'保留处理后的所有列,再调整步骤顺序或单独处理age_interval:

preprocessing = ColumnTransformer([
    ("ord", ord_pipeline, ['Pclass']),
    ("total_relatives", total_relatives_pipeline, ['SibSp', 'Parch']),
    ("age_processing", age_processor(), ['Pclass', 'Age']),
    ("cat", cat_pipeline, ['Sex', 'Embarked', 'traveling_category']),
    ("num", num_pipeline, ['Fare']),
], remainder='passthrough')

内容的提问来源于stack exchange,提问作者Silvio sjsj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 18:23:17