You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用category_encoders中OrdinalEncoder的fit_transform/transform报错求助

解决category_encoders.OrdinalEncoder在Kaggle上的报错问题

问题场景

本地Jupyter环境中使用category_encoders.OrdinalEncoder做编码运行正常,但提交到Kaggle竞赛时出现报错,两次修改参数均未解决。

最初代码

import category_encoders as ce
encoder = ce.OrdinalEncoder(cols = [['MSZoning', 'Street', 'LotShape', 'LandContour', 'Utilities', 'LotConfig', 'LandSlope', 'Neighborhood', 'Condition1', 'Condition2', 'BldgType', 'HouseStyle', 'RoofStyle', 'RoofMatl','Exterior1st', 'Exterior2nd', 'MasVnrType', 'ExterQual', 'ExterCond', 'Foundation', 'BsmtQual','BsmtCond', 'BsmtExposure', 'BsmtFinType1', 'BsmtFinType2', 'Heating', 'HeatingQC', 'CentralAir','Electrical', 'KitchenQual', 'Functional', 'GarageType', 'GarageFinish', 'GarageQual', 'GarageCond','PavedDrive', 'SaleType', 'SaleCondition']])

X_train = encoder.fit_transform(X_train)
X_test = encoder.transform(X_test)

第一次报错

TypeError                                 Traceback (most recent call last)
# # /tmp/ipykernel_27/3436411182.py in <module>
1 X_train = encoder.fit_transform(X_train)
2 X_test = encoder.transform(X_test)
# # /opt/conda/lib/python3.7/site-packages/sklearn/base.py in fit_transform(self, X, y, **fit_params)
850         if y is None:
851             # fit method of arity 1 (unsupervised transformation)
852             return self.fit(X, **fit_params).transform(X)
853         else:
854             # fit method of arity 2 (supervised transformation)
# # /opt/conda/lib/python3.7/site-packages/category_encoders/utils.py in fit(self, X, y, **kwargs)
297         self._determine_fit_columns(X)
298
299         if not set(self.cols).issubset(X.columns):
300             raise ValueError('X does not contain the columns listed in cols')
301
TypeError: unhashable type: 'list'

修改后的代码(改用单层元组)

encoder = ce.OrdinalEncoder(cols = [('MSZoning', 'Street', 'LotShape', 'LandContour', 'Utilities', 'LotConfig', 'LandSlope', 
                                    'Neighborhood', 'Condition1', 'Condition2', 'BldgType', 'HouseStyle', 'RoofStyle', 'RoofMatl',
                                    'Exterior1st', 'Exterior2nd', 'MasVnrType', 'ExterQual', 'ExterCond', 'Foundation', 'BsmtQual',
                                    'BsmtCond', 'BsmtExposure', 'BsmtFinType1', 'BsmtFinType2', 'Heating', 'HeatingQC', 'CentralAir',
                                    'Electrical', 'KitchenQual', 'Functional', 'GarageType', 'GarageFinish', 'GarageQual', 'GarageCond',
                                    'PavedDrive', 'SaleType', 'SaleCondition')])

第二次报错

ValueError                                Traceback (most recent call last)
/tmp/ipykernel_27/3436411182.py in <module>
----> 1 X_train = encoder.fit_transform(X_train)
      2 X_test = encoder.transform(X_test)

/opt/conda/lib/python3.7/site-packages/sklearn/base.py in fit_transform(self, X, y, **fit_params)
    850         if y is None:
    851             # fit method of arity 1 (unsupervised transformation)
---> 852             return self.fit(X, **fit_params).transform(X)
    853         else:
    854             # fit method of arity 2 (supervised transformation)

/opt/conda/lib/python3.7/site-packages/category_encoders/utils.py in fit(self, X, y, **kwargs)
    298 
    299         if not set(self.cols).issubset(X.columns):
---> 300  raise ValueError('X does not contain the columns listed in cols')
    301 
    302         if self.handle_missing == 'error':

ValueError: X does not contain the columns listed in cols

错误原因

  1. 第一次报错:cols参数要求传入单个列名字符串组成的可迭代对象,你传了双层列表(列表嵌套列表),代码执行到set(self.cols)时,列表是不可哈希类型,触发TypeError。
  2. 第二次报错:改用元组后,你把所有列名放在了一个元组里并套进列表,相当于让编码器去寻找一个名为这个元组的列(而非元组内的每个列),自然找不到该列,触发ValueError。

正确解决方案

将cols设置为列名字符串组成的单层列表,直接传入所有需要编码的列名即可:

import category_encoders as ce

encoder = ce.OrdinalEncoder(cols = [
    'MSZoning', 'Street', 'LotShape', 'LandContour', 'Utilities', 
    'LotConfig', 'LandSlope', 'Neighborhood', 'Condition1', 'Condition2', 
    'BldgType', 'HouseStyle', 'RoofStyle', 'RoofMatl', 'Exterior1st', 
    'Exterior2nd', 'MasVnrType', 'ExterQual', 'ExterCond', 'Foundation', 
    'BsmtQual', 'BsmtCond', 'BsmtExposure', 'BsmtFinType1', 'BsmtFinType2', 
    'Heating', 'HeatingQC', 'CentralAir', 'Electrical', 'KitchenQual', 
    'Functional', 'GarageType', 'GarageFinish', 'GarageQual', 'GarageCond', 
    'PavedDrive', 'SaleType', 'SaleCondition'
])

X_train = encoder.fit_transform(X_train)
X_test = encoder.transform(X_test)

额外注意事项

  • 本地与Kaggle运行差异可能源于category_encoders版本不同,本地版本可能对参数格式有容错,而Kaggle上的版本严格遵循参数规范。
  • 若仍提示列名不存在,可通过print(X_train.columns.tolist())输出所有列名,与cols中的列名逐一比对,确认是否存在大小写、前后空格等细微差异。

内容的提问来源于stack exchange,提问作者gaaab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 07:55:29