使用Linear Regression预测美元薪资时遇未知类别ValueError问题求助
解决测试集未知类别导致的OneHotEncoder转换错误
方案1:设置OneHotEncoder忽略未知类别
直接在OneHotEncoder中添加handle_unknown='ignore'参数,测试集里的未知类别会被对应特征列全设为0,不会触发报错。
修改后的预处理配置代码:
from sklearn.preprocessing import OneHotEncoder, StandardScaler from sklearn.compose import ColumnTransformer drop_cols = ['Currency', 'Company location', 'Salary', 'Title'] num_attributes = ['Year'] one_hot_attributes = ['Experience', 'Type', 'Remote work', 'Residence', 'Company size'] data.drop(drop_cols, 1, inplace=True) preprocessor = ColumnTransformer([ ('nums', StandardScaler(), num_attributes), # 添加handle_unknown='ignore'参数 ('one_hot', OneHotEncoder(drop='first', sparse=False, handle_unknown='ignore'), one_hot_attributes)], remainder='passthrough')
注意:此方法会直接丢弃未知类别的信息,适合未知类别样本量极少的场景。
方案2:训练阶段合并罕见类别
在训练集预处理时,将出现次数低于阈值的类别合并为Other,测试集里的未知类别也统一映射到Other,保留部分类别信息。
示例代码:
# 先处理Residence列的罕见类别(对应报错的第3列) threshold = 5 # 可根据数据集调整阈值 residence_counts = data['Residence'].value_counts() data['Residence'] = data['Residence'].apply(lambda x: 'Other' if residence_counts[x] < threshold else x) # 后续的预处理和管道代码保持不变 preprocessor = ColumnTransformer([ ('nums', StandardScaler(), num_attributes), ('one_hot', OneHotEncoder(drop='first', sparse=False), one_hot_attributes)], remainder='passthrough') pipe = Pipeline(steps =[ ('preprocessor', preprocessor), ('model', LinearRegression()), ]) pipe.fit(X_train, y_train)
方案3:自定义转换器处理未知类别
如果需要更灵活的处理逻辑(比如不同列用不同规则),可以自定义一个转换器,在预处理阶段统一处理测试集的未知类别。
示例代码:
from sklearn.base import BaseEstimator, TransformerMixin from sklearn.preprocessing import OneHotEncoder, StandardScaler from sklearn.compose import ColumnTransformer from sklearn.pipeline import Pipeline class UnknownCategoryHandler(BaseEstimator, TransformerMixin): def fit(self, X, y=None): # 记录训练集中每个列的所有类别 self.col_categories = {col: set(X[col].unique()) for col in X.columns} return self def transform(self, X): # 将测试集中的未知类别替换为Other for col in X.columns: X[col] = X[col].apply(lambda x: x if x in self.col_categories[col] else 'Other') return X # 修改管道,先处理未知类别再做编码 drop_cols = ['Currency', 'Company location', 'Salary', 'Title'] num_attributes = ['Year'] one_hot_attributes = ['Experience', 'Type', 'Remote work', 'Residence', 'Company size'] data.drop(drop_cols, 1, inplace=True) preprocessor = ColumnTransformer([ ('nums', StandardScaler(), num_attributes), ('one_hot', OneHotEncoder(drop='first', sparse=False), one_hot_attributes)], remainder='passthrough') pipe = Pipeline(steps =[ ('unknown_handler', UnknownCategoryHandler()), ('preprocessor', preprocessor), ('model', LinearRegression()), ]) pipe.fit(X_train, y_train) prediction = pipe.predict(X_test)
内容的提问来源于stack exchange,提问作者Alix Blaine
相关产品推荐
相关产品推荐

