解决OneHotEncoder初始化时出现“unexpected keyword argument 'categorical_features'”错误的问题
解决OneHotEncoder的
categorical_features参数报错问题 这个问题我之前踩过坑!你遇到的错误是因为scikit-learn版本更新后,OneHotEncoder的categorical_features参数已经被移除了——从0.22版本开始这个参数就被弃用,后续版本直接删掉了,官方推荐用更灵活的方式来指定编码列。
针对你的代码场景,给你两种简单的解决办法:
方法一:直接移除categorical_features参数(适合单列编码)
你的代码里y已经是reshape后的单列数组,OneHotEncoder默认会处理输入的所有列,所以完全不需要指定categorical_features参数,修改后的代码如下:
from sklearn.preprocessing import LabelEncoder , OneHotEncoder y=rooms_df['room type'].values y_labelencoder = LabelEncoder () y = y_labelencoder.fit_transform (y) y=y.reshape(-1,1) # 移除categorical_features参数 onehotencoder = OneHotEncoder() Y= onehotencoder.fit_transform(y) print(Y.shape)
方法二:用ColumnTransformer指定编码列(适合多列场景)
如果之后你需要处理DataFrame里的多列特征,只想对特定列做独热编码,推荐用ColumnTransformer,这是官方现在最推荐的方式,示例代码如下:
from sklearn.preprocessing import LabelEncoder, OneHotEncoder from sklearn.compose import ColumnTransformer # 直接从DataFrame处理,新版本OneHotEncoder可直接处理字符串类型,无需先转数值 ct = ColumnTransformer( transformers=[ # 指定转换器名称、编码工具、要编码的列名 ('room_type_encoder', OneHotEncoder(), ['room type']) ], # 剩余列保持原样(如果有其他特征列的话,不需要可省略此参数) remainder='passthrough' ) # 拟合转换,结果是稀疏矩阵,需要转成数组可加.toarray() Y = ct.fit_transform(rooms_df) print(Y.shape)
额外提一句:新版本的OneHotEncoder已经支持直接处理字符串类型的分类特征,不需要先用LabelEncoder转成数值,上面的代码就省略了LabelEncoder步骤,更简洁高效。
不建议你降级scikit-learn版本来适配旧参数,新版本的API设计更合理,也能避免后续更多兼容性问题。
内容的提问来源于stack exchange,提问作者taufan hidayat
相关产品推荐
相关产品推荐

