使用sklearn Pipeline训练时报ValueError: could not convert string to float错误
sklearn Pipeline使用OneHotEncoder时出现string转float错误的解决方案
错误原因
- ColumnTransformer的所有transformer是并行执行,所有transformer的输入都是原始的X_train数据,不会先跑第一个transformer再把输出传给第二个。你当前的写法里
onehot步骤拿到的workclass、education、native-country还是原始带缺失值的版本,没有经过填充;同时imputer步骤输出的3个填充后的字符串列会直接拼接到最终的预处理结果中,这3列没有经过独热编码,依然是字符串格式,模型训练时尝试转float就会触发报错。 - 你没有指定剩余数值列(比如age、hours-per-week等)的处理逻辑,如果你希望保留这些数值特征用于训练,也需要明确配置。
解决方案
把填充、独热编码的步骤串行封装到同一个子Pipeline里,再交给ColumnTransformer调用,同时配置数值列的处理逻辑,修正后的代码如下:
import json import numpy as np import pandas as pd from sklearn.model_selection import train_test_split from sklearn.preprocessing import LabelEncoder from sklearn.ensemble import RandomForestClassifier from sklearn.pipeline import Pipeline from sklearn.compose import ColumnTransformer from sklearn.preprocessing import OneHotEncoder from sklearn.impute import SimpleImputer df = pd.read_csv('https://raw.githubusercontent.com/pplonski/datasets-for-start/master/adult/data.csv', skipinitialspace=True) x_cols = [c for c in df.columns if c!='income'] X = df[x_cols] y = df['income'] y = LabelEncoder().fit_transform(y) X_train, X_test, y_train, y_test = train_test_split(X,y, test_size=0.3) # 区分离散特征和数值特征 cat_cols = ['workclass', 'education', 'marital-status', 'occupation', 'relationship', 'race', 'sex','native-country'] num_cols = [c for c in x_cols if c not in cat_cols] # 离散特征的处理流水线:先填充缺失值,再独热编码 cat_pipeline = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='most_frequent')), ('onehot', OneHotEncoder(handle_unknown='ignore')) ]) preprocessor = ColumnTransformer( transformers=[ ('cat', cat_pipeline, cat_cols), ('num', 'passthrough', num_cols) # 数值特征直接保留 ] ) clf = Pipeline([('preprocessor', preprocessor), ('classifier', RandomForestClassifier())]) clf.fit(X_train, y_train)
内容的提问来源于stack exchange,提问作者DanNg
相关产品推荐
相关产品推荐

