You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scikit-learn FunctionTransformer报错:DataFrame无str属性/形状无效

自定义FunctionTransformer在Sklearn Pipeline中报错的原因与解决方法

核心错误原因

所有报错的本质都是Sklearn转换器(包括FunctionTransformer)默认接收/输出二维结构(如DataFrame或2D数组),但你的自定义函数是针对一维Series编写的,导致类型不匹配:

  1. 当通过Pipeline传递单列数据(如['bathrooms_text'])时,FunctionTransformer接收的是单列DataFrame,而非Series,调用.str属性会直接报错(DataFrame没有.str方法)。
  2. 若强行用squeeze()转成Series,返回时如果是一维结构,后续Sklearn步骤(如SimpleImputer)会因形状不匹配报错(要求输入为(n_samples, n_features)的二维格式)。
  3. price列的问题是自定义函数未彻底将所有字符串转为数值,或返回结构不符合后续步骤要求,导致SimpleImputer仍遇到字符串类型数据。

修正方案

1. 修改自定义函数,适配二维输入+输出

所有自定义函数需要:

  • 先将输入的单列DataFrame转为Series处理
  • 处理完成后返回二维结构(单列DataFrame或2D数组),适配Sklearn的输入要求

修正后的clean_bathrooms函数

import numpy as np
import pandas as pd

def clean_bathrooms(bathrooms_df):
    # 将单列DataFrame转为Series
    bathrooms_text = bathrooms_df.squeeze()
    bathrooms_text = bathrooms_text.copy()
    
    pattern = r'(\d.?\d?)\s'
    pattern2 = r'(Half)'
    
    # 处理Half浴室
    bathrooms_text.loc[bathrooms_text.str.contains(pattern2, na=False)] = 0.5
    # 提取数字部分
    bathrooms_text = bathrooms_text.str.extract(pattern)
    # 返回二维数组(适配Sklearn后续步骤的输入格式)
    return bathrooms_text.astype(float).values.reshape(-1, 1)

修正后的clean_property_type函数

def clean_property_type(property_type_df):
    # 将单列DataFrame转为Series
    property_type_col = property_type_df.squeeze()
    property_type_col = property_type_col.copy()
    
    # 合并类别
    property_type_col.loc[property_type_col.str.contains(r'Entire|Tiny home', na=False)] = 'Entire Unit'
    property_type_col.loc[property_type_col.str.contains(r'[Rr]oom', na=False)] = 'Single Room'
    property_type_col.loc[property_type_col.str.contains(r'Camp', na=False)] = 'Camping'
    
    # 标记其他类别为NaN
    property_type_col.loc[~property_type_col.isin(['Camping', 'Single Room', 'Entire Unit'])] = np.nan
    # 返回单列DataFrame(保持二维结构)
    return property_type_col.to_frame()

修正后的extract_price函数

def extract_price(price_df):
    # 将单列DataFrame转为Series
    price_series = price_df.squeeze()
    # 批量替换$和逗号,避免逐行apply的低效
    price_series = price_series.str.replace('$', '', regex=False).str.replace(',', '', regex=False)
    # 转为float并返回单列DataFrame
    return price_series.astype(float).to_frame()

2. 保持Pipeline结构不变,直接使用修正后的函数

无需修改Pipeline代码,直接将修正后的函数传入FunctionTransformer即可:

修正后的price列Pipeline示例

from sklearn.preprocessing import FunctionTransformer, StandardScaler
from sklearn.impute import SimpleImputer
from sklearn.pipeline import make_pipeline

clean_price = FunctionTransformer(func=extract_price)
clean_price_pipeline = make_pipeline(
    clean_price,
    SimpleImputer(strategy='mean'),
    FunctionTransformer(np.log1p),
    StandardScaler()
)

# 测试运行
t = df[['price']].copy()
clean_price_pipeline.fit_transform(t)

额外优化:简化FunctionTransformer的使用

如果不想修改自定义函数,也可以在FunctionTransformer中直接处理输入输出的形状转换,例如:

bathrooms_pipeline = make_pipeline(
    FunctionTransformer(func=lambda x: clean_bathrooms(x.squeeze()).values.reshape(-1,1)),
    SimpleImputer(strategy='mean'),
    StandardScaler()
)

但这种方式可读性较差,建议优先修改自定义函数适配Sklearn的结构要求。


内容的提问来源于stack exchange,提问作者neal301

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 05:50:55