You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用OrdinalEncoder触发Shape mismatch错误,请求排查代码问题

问题分析与修复

错误根源

触发ValueError: Shape mismatch: if categories is an array, it has to be of shape (n_features,).的原因是**OrdinalEncoder的categories参数格式错误**:

  • 手动指定类别时,categories必须是二维数组,每个子数组对应一个待编码特征的类别顺序
  • 你当前传入的["Strong","Mild"]是一维数组,而要处理的是1个特征(原数据第3列),不符合(n_features,)的形状要求(当n_features=1时,需要外层嵌套一个列表)

修复后的代码

from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder, StandardScaler
from sklearn.pipeline import Pipeline

trf1 = ColumnTransformer([("Infuse_val", SimpleImputer(strategy="mean"), [0])], remainder="passthrough")
trf4 = ColumnTransformer([("One_hot", OneHotEncoder(sparse=False, handle_unknown="ignore"), [1,4])], remainder="passthrough")
# 修正categories为二维数组,匹配1个特征的要求
trf2 = ColumnTransformer([("Ord_encode", OrdinalEncoder(categories=[["Strong","Mild"]]), [3])], remainder="passthrough")
trf3 = ColumnTransformer([("scale", StandardScaler(), [0,2])], remainder="passthrough")

pipe = Pipeline([
    ('trf1', trf1),
    ('trf2', trf2),
    ('trf3', trf3),
    ('trf4', trf4),
])
# 注意:原代码中的y_tarin是拼写错误,需改为y_train
pipe.fit(x_train, y_train)

额外提示

  1. 原代码中y_tarin是笔误,必须修正为y_train,否则会触发NameError
  2. 预处理流程的索引是正确的:由于每一步ColumnTransformer使用remainder="passthrough",输出特征顺序为「处理后的列 + 未处理的原列」,后续步骤指定的索引(如trf2的[3]、trf3的[0,2])能正确匹配原数据列位置,无需调整

内容的提问来源于stack exchange,提问作者Akash Mukherjee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 19:45:55