RandomForestClassifier特征数不匹配ValueError问题求助
问题描述
ValueError: X has 42 features, but RandomForestClassifier is expecting 1998 features as input.
本人是机器学习与Python新手,正在用RandomForestClassifier构建模型预测计划变更失败概率。训练阶段代码运行正常,训练集特征数为1998,但预测阶段数据处理后仅42个特征,触发上述特征数不匹配错误。相关代码及训练集维度如下:
训练与预测代码
# load the dataset df = pd.read_csv("training v.3.csv", encoding='cp1252') # Keep the neccessary columns that's needed cols_to_keep = [ "close_code", "assignment_group", "assignment_group_category", "cmdb_ci", "type", "sys_created_on_day_of_week", "cab_required", "cmdb_ci.environment", "size_backout", ] df = df[cols_to_keep] # drop rows with missing values df = df.dropna() # Define features & labels feat = df.drop(columns=["close_code"]) label = df["close_code"] # one-hot encoding of the categotical columns feature = pd.get_dummies(feat, columns=feat.columns) # initialize the one-hot encoder encoder = OneHotEncoder(sparse_output=False, handle_unknown="ignore") # fit the hot encoder to the categorical variables encoder.fit(feature) # Transform the categorical variables X_train_encoded = encoder.transform(feature) # Scale the data scaler = StandardScaler() feature_encoded = scaler.fit_transform(X_train_encoded) # Split the data into a training set and a test set X_train, X_test, y_train, y_test = train_test_split( feature_encoded, label, train_size=0.65, test_size=0.35, random_state=42 ) from collections import Counter # Get the class counts counts = Counter(y_train) # Determine the minority class minority_class_label = min(counts, key=counts.get) # Define sampling_strategy sampling_strategy = {minority_class_label: int(counts[minority_class_label])} from imblearn.combine import SMOTETomek # Apply SMOTE-TOMEK smt = SMOTETomek(sampling_strategy=sampling_strategy) X_train_resampled, y_train_resampled = smt.fit_resample(X_train, y_train) # Train the Random Forest Classifier on the oversampled data rfc = RandomForestClassifier( n_estimators=1600, min_samples_split=6, min_samples_leaf=1, max_depth=15, class_weight="balanced", ) # train the classifier on the oversampled train data rfc.fit(X_train_resampled, y_train_resampled) # Evaluate the classifier on the test set y_pred = rfc.predict(X_test) # print the feature importances for name, importance in zip(feat, rfc.feature_importances_): print(name, "=", importance) # # print out the classification report print(classification_report(y_test, y_pred)) import joblib from sklearn import preprocessing rfc_loaded = RandomForestClassifier() X_train_encoded = X_train_encoded[:len(y_train_resampled), :] rfc_loaded.fit(X_train_encoded, y_train_resampled) d = preprocessing.LabelEncoder i = [d, rfc_loaded] filename = 'rfc_model.model' joblib.dump(i, filename) prediction_df = pd.read_csv("prediction.csv", encoding='cp1252') # Load the model d, rfc_loaded = joblib.load(filename) # Drop rows with missing values # prediction_df = prediction_df.dropna cols_to_keep = [ "assignment_group", "assignment_group.u_accenture_category", "cmdb_ci", "type", "sys_created_on_day_of_week", "cab_required", "cmdb_ci.environment", "size_backout" ] prediction_df=prediction_df[cols_to_keep] prediction_deep = prediction_df.copy(deep=True) # initialize the one-hot encoder encoder = OneHotEncoder(sparse_output=False, handle_unknown="ignore") # fit the hot encoder to the categorical variables encoder.fit(prediction_df) # Transform the categorical variables X_train_encoded = encoder.transform(prediction_df) # Read Target Column prediction_df['target'] = 0 # Set Numerical Columns numerical_columns = [col for col in prediction_df.columns if (prediction_df[col].dtype == 'int64' or prediction_df[col].dtype == 'float64') and col != 'target'] # One-Hot Encode the categorical variables prediction_df = pd.get_dummies(prediction_df) # Scale the data scaler = StandardScaler() prediction_df = scaler.fit_transform(prediction_df) filename = 'rfc_model.model' # Load the model d, rfc_loaded = joblib.load(filename) # Make Predictions prediction_df['predict'] = pd.DataFrame(rfc_loaded.predict(prediction_df)) prediction_df['probability'] = pd.DataFrame(rfc_loaded.predict_proba(prediction_df) [:, 1]) count=sum(prediction_df['predict'] ==1 ) print("Check" , count) prediction_df.to_csv(r"C:\Users\shaina.bowser\OneDrive - Accenture\Desktop\ProjectLemonade\results.csv")
训练集维度
X_train.shape (3515, 1998) X_train_resample.shape (3487, 1998) X_test.shape (1894, 1998)
核心问题排查
- 特征处理流程完全不一致:训练时先对特征列做
pd.get_dummies,再用OneHotEncoder拟合训练集特征后转换,最后标准化;但预测阶段不仅重新拟合了新的OneHotEncoder,还额外添加了target列、重复用pd.get_dummies,导致特征结构完全偏离训练集。 - 未保存关键预处理组件:训练时没有保存拟合好的
OneHotEncoder和StandardScaler,反而保存了未实例化的LabelEncoder类,预测时无法复用训练集的特征编码规则,只能基于小样本预测数据生成少量特征。 - 特征列名不匹配:训练时保留的特征是
assignment_group_category,但预测时用的是assignment_group.u_accenture_category,直接导致缺失训练集的关键特征维度。 - 冗余的模型训练步骤:训练最后额外训练了
rfc_loaded,用截断后的X_train_encoded,完全没必要且可能导致模型参数异常。
修正后的代码示例
训练阶段修正代码
import pandas as pd from sklearn.model_selection import train_test_split from sklearn.preprocessing import OneHotEncoder, StandardScaler from sklearn.ensemble import RandomForestClassifier from sklearn.metrics import classification_report from imblearn.combine import SMOTETomek from collections import Counter import joblib # 加载数据集 df = pd.read_csv("training v.3.csv", encoding='cp1252') # 保留需要的特征列(列名与后续预测严格统一) cols_to_keep = [ "close_code", "assignment_group", "assignment_group_category", "cmdb_ci", "type", "sys_created_on_day_of_week", "cab_required", "cmdb_ci.environment", "size_backout", ] df = df[cols_to_keep] df = df.dropna() # 划分特征和标签 feat = df.drop(columns=["close_code"]) label = df["close_code"] # 初始化并拟合OneHotEncoder(直接处理原始分类特征,避免冗余编码) encoder = OneHotEncoder(sparse_output=False, handle_unknown="ignore") X_encoded = encoder.fit_transform(feat) # 标准化 scaler = StandardScaler() X_scaled = scaler.fit_transform(X_encoded) # 划分训练集和测试集 X_train, X_test, y_train, y_test = train_test_split( X_scaled, label, train_size=0.65, test_size=0.35, random_state=42 ) # 类别平衡(若需调整采样策略,可修改sampling_strategy的值) counts = Counter(y_train) minority_class_label = min(counts, key=counts.get) sampling_strategy = {minority_class_label: counts[minority_class_label]} smt = SMOTETomek(sampling_strategy=sampling_strategy) X_train_resampled, y_train_resampled = smt.fit_resample(X_train, y_train) # 训练模型 rfc = RandomForestClassifier( n_estimators=1600, min_samples_split=6, min_samples_leaf=1, max_depth=15, class_weight="balanced", random_state=42 ) rfc.fit(X_train_resampled, y_train_resampled) # 评估模型 y_pred = rfc.predict(X_test) print(classification_report(y_test, y_pred)) # 保存模型+预处理组件(核心:复用训练时的编码/标准化规则) joblib.dump({"encoder": encoder, "scaler": scaler, "model": rfc}, "rfc_model_with_preprocessors.joblib")
预测阶段修正代码
import pandas as pd import joblib # 加载预测数据 prediction_df = pd.read_csv("prediction.csv", encoding='cp1252') # 加载保存的模型和预处理组件 saved_objects = joblib.load("rfc_model_with_preprocessors.joblib") encoder = saved_objects["encoder"] scaler = saved_objects["scaler"] rfc = saved_objects["model"] # 保留与训练集完全一致的特征列 cols_to_keep = [ "assignment_group", "assignment_group_category", "cmdb_ci", "type", "sys_created_on_day_of_week", "cab_required", "cmdb_ci.environment", "size_backout" ] prediction_df = prediction_df[cols_to_keep] prediction_df = prediction_df.dropna() # 处理缺失值 # 复用训练好的编码器和标准化器:绝对不能重新fit! X_pred_encoded = encoder.transform(prediction_df) X_pred_scaled = scaler.transform(X_pred_encoded) # 执行预测 predictions = rfc.predict(X_pred_scaled) probabilities = rfc.predict_proba(X_pred_scaled)[:, 1] # 整理结果 prediction_df["predict"] = predictions prediction_df["probability"] = probabilities print("预测为1的数量:", sum(prediction_df["predict"] == 1)) # 保存结果 prediction_df.to_csv(r"C:\Users\shaina.bowser\OneDrive - Accenture\Desktop\ProjectLemonade\results.csv", index=False)
内容的提问来源于stack exchange,提问作者Shaina Bowser
相关产品推荐
相关产品推荐

