You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Linear Regression预测美元薪资时遇未知类别ValueError问题求助

解决测试集未知类别导致的OneHotEncoder转换错误

方案1:设置OneHotEncoder忽略未知类别

直接在OneHotEncoder中添加handle_unknown='ignore'参数,测试集里的未知类别会被对应特征列全设为0,不会触发报错。

修改后的预处理配置代码:

from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.compose import ColumnTransformer

drop_cols = ['Currency', 'Company location', 'Salary', 'Title']
num_attributes = ['Year']
one_hot_attributes = ['Experience', 'Type', 'Remote work', 'Residence', 'Company size']

data.drop(drop_cols, 1, inplace=True)

preprocessor = ColumnTransformer([
    ('nums', StandardScaler(), num_attributes),
    # 添加handle_unknown='ignore'参数
    ('one_hot', OneHotEncoder(drop='first', sparse=False, handle_unknown='ignore'), one_hot_attributes)], 
                                remainder='passthrough')

注意:此方法会直接丢弃未知类别的信息,适合未知类别样本量极少的场景。

方案2:训练阶段合并罕见类别

在训练集预处理时,将出现次数低于阈值的类别合并为Other,测试集里的未知类别也统一映射到Other,保留部分类别信息。

示例代码:

# 先处理Residence列的罕见类别(对应报错的第3列)
threshold = 5  # 可根据数据集调整阈值
residence_counts = data['Residence'].value_counts()
data['Residence'] = data['Residence'].apply(lambda x: 'Other' if residence_counts[x] < threshold else x)

# 后续的预处理和管道代码保持不变
preprocessor = ColumnTransformer([
    ('nums', StandardScaler(), num_attributes),
    ('one_hot', OneHotEncoder(drop='first', sparse=False), one_hot_attributes)], 
                                remainder='passthrough')

pipe = Pipeline(steps =[
    ('preprocessor', preprocessor),
    ('model', LinearRegression()),
])
pipe.fit(X_train, y_train)

方案3:自定义转换器处理未知类别

如果需要更灵活的处理逻辑(比如不同列用不同规则),可以自定义一个转换器,在预处理阶段统一处理测试集的未知类别。

示例代码:

from sklearn.base import BaseEstimator, TransformerMixin
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline

class UnknownCategoryHandler(BaseEstimator, TransformerMixin):
    def fit(self, X, y=None):
        # 记录训练集中每个列的所有类别
        self.col_categories = {col: set(X[col].unique()) for col in X.columns}
        return self
    
    def transform(self, X):
        # 将测试集中的未知类别替换为Other
        for col in X.columns:
            X[col] = X[col].apply(lambda x: x if x in self.col_categories[col] else 'Other')
        return X

# 修改管道,先处理未知类别再做编码
drop_cols = ['Currency', 'Company location', 'Salary', 'Title']
num_attributes = ['Year']
one_hot_attributes = ['Experience', 'Type', 'Remote work', 'Residence', 'Company size']

data.drop(drop_cols, 1, inplace=True)

preprocessor = ColumnTransformer([
    ('nums', StandardScaler(), num_attributes),
    ('one_hot', OneHotEncoder(drop='first', sparse=False), one_hot_attributes)], 
                                remainder='passthrough')

pipe = Pipeline(steps =[
    ('unknown_handler', UnknownCategoryHandler()),
    ('preprocessor', preprocessor),
    ('model', LinearRegression()),
])

pipe.fit(X_train, y_train)
prediction = pipe.predict(X_test)

内容的提问来源于stack exchange,提问作者Alix Blaine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 04:30:53