Scikit-learn FunctionTransformer报错:DataFrame无str属性/形状无效
自定义FunctionTransformer在Sklearn Pipeline中报错的原因与解决方法
核心错误原因
所有报错的本质都是Sklearn转换器(包括FunctionTransformer)默认接收/输出二维结构(如DataFrame或2D数组),但你的自定义函数是针对一维Series编写的,导致类型不匹配:
- 当通过Pipeline传递单列数据(如
['bathrooms_text'])时,FunctionTransformer接收的是单列DataFrame,而非Series,调用.str属性会直接报错(DataFrame没有.str方法)。 - 若强行用
squeeze()转成Series,返回时如果是一维结构,后续Sklearn步骤(如SimpleImputer)会因形状不匹配报错(要求输入为(n_samples, n_features)的二维格式)。 - price列的问题是自定义函数未彻底将所有字符串转为数值,或返回结构不符合后续步骤要求,导致SimpleImputer仍遇到字符串类型数据。
修正方案
1. 修改自定义函数,适配二维输入+输出
所有自定义函数需要:
- 先将输入的单列DataFrame转为Series处理
- 处理完成后返回二维结构(单列DataFrame或2D数组),适配Sklearn的输入要求
修正后的clean_bathrooms函数
import numpy as np import pandas as pd def clean_bathrooms(bathrooms_df): # 将单列DataFrame转为Series bathrooms_text = bathrooms_df.squeeze() bathrooms_text = bathrooms_text.copy() pattern = r'(\d.?\d?)\s' pattern2 = r'(Half)' # 处理Half浴室 bathrooms_text.loc[bathrooms_text.str.contains(pattern2, na=False)] = 0.5 # 提取数字部分 bathrooms_text = bathrooms_text.str.extract(pattern) # 返回二维数组(适配Sklearn后续步骤的输入格式) return bathrooms_text.astype(float).values.reshape(-1, 1)
修正后的clean_property_type函数
def clean_property_type(property_type_df): # 将单列DataFrame转为Series property_type_col = property_type_df.squeeze() property_type_col = property_type_col.copy() # 合并类别 property_type_col.loc[property_type_col.str.contains(r'Entire|Tiny home', na=False)] = 'Entire Unit' property_type_col.loc[property_type_col.str.contains(r'[Rr]oom', na=False)] = 'Single Room' property_type_col.loc[property_type_col.str.contains(r'Camp', na=False)] = 'Camping' # 标记其他类别为NaN property_type_col.loc[~property_type_col.isin(['Camping', 'Single Room', 'Entire Unit'])] = np.nan # 返回单列DataFrame(保持二维结构) return property_type_col.to_frame()
修正后的extract_price函数
def extract_price(price_df): # 将单列DataFrame转为Series price_series = price_df.squeeze() # 批量替换$和逗号,避免逐行apply的低效 price_series = price_series.str.replace('$', '', regex=False).str.replace(',', '', regex=False) # 转为float并返回单列DataFrame return price_series.astype(float).to_frame()
2. 保持Pipeline结构不变,直接使用修正后的函数
无需修改Pipeline代码,直接将修正后的函数传入FunctionTransformer即可:
修正后的price列Pipeline示例
from sklearn.preprocessing import FunctionTransformer, StandardScaler from sklearn.impute import SimpleImputer from sklearn.pipeline import make_pipeline clean_price = FunctionTransformer(func=extract_price) clean_price_pipeline = make_pipeline( clean_price, SimpleImputer(strategy='mean'), FunctionTransformer(np.log1p), StandardScaler() ) # 测试运行 t = df[['price']].copy() clean_price_pipeline.fit_transform(t)
额外优化:简化FunctionTransformer的使用
如果不想修改自定义函数,也可以在FunctionTransformer中直接处理输入输出的形状转换,例如:
bathrooms_pipeline = make_pipeline( FunctionTransformer(func=lambda x: clean_bathrooms(x.squeeze()).values.reshape(-1,1)), SimpleImputer(strategy='mean'), StandardScaler() )
但这种方式可读性较差,建议优先修改自定义函数适配Sklearn的结构要求。
内容的提问来源于stack exchange,提问作者neal301
相关产品推荐
相关产品推荐

