You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

RandomForestClassifier特征数不匹配ValueError问题求助

问题描述

ValueError: X has 42 features, but RandomForestClassifier is expecting 1998 features as input.

本人是机器学习与Python新手,正在用RandomForestClassifier构建模型预测计划变更失败概率。训练阶段代码运行正常,训练集特征数为1998,但预测阶段数据处理后仅42个特征,触发上述特征数不匹配错误。相关代码及训练集维度如下:

训练与预测代码

# load the dataset 
df = pd.read_csv("training v.3.csv", encoding='cp1252')

# Keep the neccessary columns that's needed
cols_to_keep = [
    "close_code",
    "assignment_group",
    "assignment_group_category",
    "cmdb_ci",
    "type",
    "sys_created_on_day_of_week",
    "cab_required",
    "cmdb_ci.environment",
    "size_backout",
]

df = df[cols_to_keep]

# drop rows with missing values
df = df.dropna()

# Define features & labels
feat = df.drop(columns=["close_code"])
label = df["close_code"]

# one-hot encoding of the categotical columns
feature = pd.get_dummies(feat, columns=feat.columns)


# initialize the one-hot encoder
encoder = OneHotEncoder(sparse_output=False, handle_unknown="ignore")

# fit the hot encoder to the categorical variables
encoder.fit(feature)

# Transform the categorical variables
X_train_encoded = encoder.transform(feature)

# Scale the data
scaler = StandardScaler()
feature_encoded = scaler.fit_transform(X_train_encoded)

# Split the data into a training set and a test set
X_train, X_test, y_train, y_test = train_test_split(
    feature_encoded, label, train_size=0.65, test_size=0.35, random_state=42
)


from collections import Counter

# Get the class counts
counts = Counter(y_train)

# Determine the minority class
minority_class_label = min(counts, key=counts.get)

# Define sampling_strategy
sampling_strategy = {minority_class_label: int(counts[minority_class_label])}

from imblearn.combine import SMOTETomek

# Apply SMOTE-TOMEK
smt = SMOTETomek(sampling_strategy=sampling_strategy)
X_train_resampled, y_train_resampled = smt.fit_resample(X_train, y_train)


# Train the Random Forest Classifier on the oversampled data
rfc = RandomForestClassifier(
    n_estimators=1600,
    min_samples_split=6,
    min_samples_leaf=1,
    max_depth=15,
    class_weight="balanced",
)

# train the classifier on the oversampled train data
rfc.fit(X_train_resampled, y_train_resampled)

# Evaluate the classifier on the test set
y_pred = rfc.predict(X_test)

# print the feature importances
for name, importance in zip(feat, rfc.feature_importances_):
    print(name, "=", importance)

# # print out the classification report
print(classification_report(y_test, y_pred))

import joblib
from sklearn import preprocessing

rfc_loaded = RandomForestClassifier()

X_train_encoded = X_train_encoded[:len(y_train_resampled), :]
rfc_loaded.fit(X_train_encoded, y_train_resampled)

d = preprocessing.LabelEncoder

i = [d, rfc_loaded]

filename = 'rfc_model.model'
joblib.dump(i, filename)

prediction_df = pd.read_csv("prediction.csv", encoding='cp1252')

# Load the model
d, rfc_loaded = joblib.load(filename)



# Drop rows with missing values
# prediction_df = prediction_df.dropna

cols_to_keep = [
    "assignment_group",
    "assignment_group.u_accenture_category",
    "cmdb_ci",
    "type",
    "sys_created_on_day_of_week",
    "cab_required",
    "cmdb_ci.environment",
    "size_backout"
]

prediction_df=prediction_df[cols_to_keep]

prediction_deep = prediction_df.copy(deep=True)

# initialize the one-hot encoder
encoder = OneHotEncoder(sparse_output=False, handle_unknown="ignore")

# fit the hot encoder to the categorical variables
encoder.fit(prediction_df)

# Transform the categorical variables
X_train_encoded = encoder.transform(prediction_df)


# Read Target Column
prediction_df['target'] = 0

# Set Numerical Columns
numerical_columns = [col for col in prediction_df.columns if                     
                    (prediction_df[col].dtype == 'int64' or prediction_df[col].dtype == 'float64') and col != 'target']

# One-Hot Encode the categorical variables
prediction_df = pd.get_dummies(prediction_df)

# Scale the data
scaler = StandardScaler()
prediction_df = scaler.fit_transform(prediction_df)


filename = 'rfc_model.model'
# Load the model
d, rfc_loaded = joblib.load(filename)

# Make Predictions
prediction_df['predict'] = pd.DataFrame(rfc_loaded.predict(prediction_df))

prediction_df['probability'] = pd.DataFrame(rfc_loaded.predict_proba(prediction_df) [:, 1])

count=sum(prediction_df['predict'] ==1 )

print("Check" , count)

prediction_df.to_csv(r"C:\Users\shaina.bowser\OneDrive - Accenture\Desktop\ProjectLemonade\results.csv")

训练集维度

X_train.shape (3515, 1998)
X_train_resample.shape (3487, 1998)
X_test.shape (1894, 1998)
核心问题排查
  • 特征处理流程完全不一致:训练时先对特征列做pd.get_dummies,再用OneHotEncoder拟合训练集特征后转换,最后标准化;但预测阶段不仅重新拟合了新的OneHotEncoder,还额外添加了target列、重复用pd.get_dummies,导致特征结构完全偏离训练集。
  • 未保存关键预处理组件:训练时没有保存拟合好的OneHotEncoder和StandardScaler,反而保存了未实例化的LabelEncoder类,预测时无法复用训练集的特征编码规则,只能基于小样本预测数据生成少量特征。
  • 特征列名不匹配:训练时保留的特征是assignment_group_category,但预测时用的是assignment_group.u_accenture_category,直接导致缺失训练集的关键特征维度。
  • 冗余的模型训练步骤:训练最后额外训练了rfc_loaded,用截断后的X_train_encoded,完全没必要且可能导致模型参数异常。
修正后的代码示例

训练阶段修正代码

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report
from imblearn.combine import SMOTETomek
from collections import Counter
import joblib

# 加载数据集
df = pd.read_csv("training v.3.csv", encoding='cp1252')

# 保留需要的特征列(列名与后续预测严格统一)
cols_to_keep = [
    "close_code",
    "assignment_group",
    "assignment_group_category",
    "cmdb_ci",
    "type",
    "sys_created_on_day_of_week",
    "cab_required",
    "cmdb_ci.environment",
    "size_backout",
]

df = df[cols_to_keep]
df = df.dropna()

# 划分特征和标签
feat = df.drop(columns=["close_code"])
label = df["close_code"]

# 初始化并拟合OneHotEncoder(直接处理原始分类特征,避免冗余编码)
encoder = OneHotEncoder(sparse_output=False, handle_unknown="ignore")
X_encoded = encoder.fit_transform(feat)

# 标准化
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X_encoded)

# 划分训练集和测试集
X_train, X_test, y_train, y_test = train_test_split(
    X_scaled, label, train_size=0.65, test_size=0.35, random_state=42
)

# 类别平衡(若需调整采样策略,可修改sampling_strategy的值)
counts = Counter(y_train)
minority_class_label = min(counts, key=counts.get)
sampling_strategy = {minority_class_label: counts[minority_class_label]}
smt = SMOTETomek(sampling_strategy=sampling_strategy)
X_train_resampled, y_train_resampled = smt.fit_resample(X_train, y_train)

# 训练模型
rfc = RandomForestClassifier(
    n_estimators=1600,
    min_samples_split=6,
    min_samples_leaf=1,
    max_depth=15,
    class_weight="balanced",
    random_state=42
)
rfc.fit(X_train_resampled, y_train_resampled)

# 评估模型
y_pred = rfc.predict(X_test)
print(classification_report(y_test, y_pred))

# 保存模型+预处理组件(核心:复用训练时的编码/标准化规则)
joblib.dump({"encoder": encoder, "scaler": scaler, "model": rfc}, "rfc_model_with_preprocessors.joblib")

预测阶段修正代码

import pandas as pd
import joblib

# 加载预测数据
prediction_df = pd.read_csv("prediction.csv", encoding='cp1252')

# 加载保存的模型和预处理组件
saved_objects = joblib.load("rfc_model_with_preprocessors.joblib")
encoder = saved_objects["encoder"]
scaler = saved_objects["scaler"]
rfc = saved_objects["model"]

# 保留与训练集完全一致的特征列
cols_to_keep = [
    "assignment_group",
    "assignment_group_category",
    "cmdb_ci",
    "type",
    "sys_created_on_day_of_week",
    "cab_required",
    "cmdb_ci.environment",
    "size_backout"
]

prediction_df = prediction_df[cols_to_keep]
prediction_df = prediction_df.dropna()  # 处理缺失值

# 复用训练好的编码器和标准化器:绝对不能重新fit!
X_pred_encoded = encoder.transform(prediction_df)
X_pred_scaled = scaler.transform(X_pred_encoded)

# 执行预测
predictions = rfc.predict(X_pred_scaled)
probabilities = rfc.predict_proba(X_pred_scaled)[:, 1]

# 整理结果
prediction_df["predict"] = predictions
prediction_df["probability"] = probabilities

print("预测为1的数量:", sum(prediction_df["predict"] == 1))

# 保存结果
prediction_df.to_csv(r"C:\Users\shaina.bowser\OneDrive - Accenture\Desktop\ProjectLemonade\results.csv", index=False)

内容的提问来源于stack exchange,提问作者Shaina Bowser

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 22:25:30