You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

蘑菇分类项目中KNeighborsClassifier预测阶段特征数量不匹配报错的解决方案咨询

Fix for "X has 61 features, but KNeighborsClassifier is expecting 74 features" Error

Hey there, I see exactly what's causing that feature mismatch issue—let's break it down and fix it step by step.

The Root Cause

You're calling fit_transform() on your test feature data (test_feature_encode = cat_encoder.fit_transform(test_feature).toarray()), which tells the OneHotEncoder to re-learn the categorical mappings from scratch using only the test data. This creates a different set of features than the ones your KNN model was trained on (since the test data might have fewer unique categories than the training set).

Instead, you should use the encoder you already fit on the training data to transform the test data—this ensures the feature dimensions match perfectly.

Corrected Code Changes

Here's the fix for the encoding step:

# Keep this line as is (fit on training data)
train_feature_encode = cat_encoder.fit_transform(train_feature).toarray()

# Change fit_transform to transform for test data
test_feature_encode = cat_encoder.transform(test_feature).toarray()

Additional Checks to Ensure Consistency

While we're at it, double-check this line to make sure you're dropping the correct column from the impute data:

# Original line: uses model_data_df.columns[11] for impute_data_df
test_feature = impute_data_df.drop(impute_data_df.columns[11], axis=1)

Using impute_data_df.columns[11] instead of model_data_df.columns[11] ensures you're targeting the right column in the imputation subset (though in most cases these will be the same, it's a safe guard against any accidental column shifts).

Why This Works

  • When you fit() the encoder on train_feature, it stores all the categorical categories present in the training data.
  • Using transform() on test_feature applies that pre-learned mapping to the test data, even if some categories are missing from the test set (your handle_unknown='ignore' parameter will handle those cases gracefully).

This way, both your training and test feature matrices will have the same number of features, and your KNN model will be able to make predictions without throwing that dimension mismatch error.

内容的提问来源于stack exchange,提问作者pevdr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:37:44