Flask部署MultinomialNB模型时特征数不匹配错误的解决方法
解决Flask部署机器学习模型时的特征数不匹配错误
问题概述
本地训练MultinomialNB模型一切正常,但通过Flask API调用预测时,出现错误:
ValueError: X has 25 features, but MultinomialNB is expecting 26 features as input
训练代码(model.py)和部署代码(app.py)如下:
model.py 代码
import pandas as pd import numpy as np import matplotlib.pyplot as plt import seaborn as sns from sklearn.preprocessing import MinMaxScaler, LabelEncoder from sklearn.model_selection import train_test_split from sklearn.preprocessing import OrdinalEncoder from sklearn.metrics import confusion_matrix , classification_report import pickle train = pd.read_csv("train_data_evaluation_part_2.csv", index_col=False) test = pd.read_csv("test_data_evaluation_part2.csv", index_col=False) df = pd.concat([train,test],axis=0) def unique_obj_col_value(df): for column in df: if df[column].dtype == 'object': print(f'{column}: {df[column].unique()}') unique_obj_col_value(df) df = df.drop('Nationality', axis=1) df['BookingsCheckedIn'] = df['BookingsCheckedIn'].replace([ 3, 1, 9, 2, 11, 12, 7, 8, 5, 6, 4, 66, 15, 29, 25, 10, 17, 13, 26, 23, 57, 40, 18, 14, 24, 19, 20, 34], 1) cols = ['DistributionChannel', 'MarketSegment'] le = LabelEncoder() for col in cols: df[col] = le.fit_transform(df[col]) pickle.dump(le, open('transform.pkl', 'wb')) df.drop(['Unnamed: 0', 'ID'], axis=1, inplace=True) df['Age'].fillna(np.mean(df['Age']),inplace=True) X = df.drop('BookingsCheckedIn', axis=1) y = df['BookingsCheckedIn'] cols_to_scale = ['Age','DaysSinceCreation', 'AverageLeadTime', 'LodgingRevenue', 'OtherRevenue','DaysSinceLastStay', 'DaysSinceFirstStay'] scaler = MinMaxScaler() X[cols_to_scale] = scaler.fit_transform(X[cols_to_scale]) print(X.shape) X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) print(X_train.shape) from sklearn.naive_bayes import MultinomialNB clf = MultinomialNB() clf.fit(X_train,y_train) clf.score(X_test,y_test) yp = clf.predict(X_test) yp[:5] y_pred = [] for element in yp: if element > 0.5: y_pred.append(1) else: y_pred.append(0) print(classification_report(y_test,y_pred)) pickle.dump(clf, open('model1.pkl', 'wb')) model = pickle.load(open('model1.pkl', 'rb'))
app.py 代码
import numpy as np from flask import Flask, request, jsonify, render_template import pickle app = Flask(__name__) model = pickle.load(open('model1.pkl','rb')) trans = pickle.load(open('transform.pkl', 'rb')) @app.route('/') def home(): return render_template('index.html') @app.route('/predict',methods=['POST']) def predict(): ''' For rendering results on HTML GUI ''' features = [x for x in request.form.values()] final_features = [np.array(features, dtype=float)] prediction = model.predict(final_features) output = round(prediction[0], 2) return render_template('index.html', prediction_text='Booked: ${}'.format(output)) if __name__ == "__main__": app.run(debug=True)
错误原因
- 预处理不完整且未复用:训练时对分类列做了LabelEncoder转换、数值列做了MinMaxScaler缩放,但部署时仅加载了LabelEncoder且未正确应用,也未加载Scaler做缩放,导致输入特征不符合模型要求。
- LabelEncoder使用错误:用同一个LabelEncoder连续处理两个分类列,最终保存的编码器仅适配最后一个列(MarketSegment),无法正确转换第一个列(DistributionChannel)。
- 输入特征数量/顺序不匹配:网页表单输入的特征数量可能少于训练时X的26列,或特征顺序与训练时不一致,导致模型接收的特征数不足。
修复步骤
步骤1:修正训练代码,保存所有必要的预处理组件
修改model.py中的预处理部分,单独保存每个分类列的编码器和数值缩放器,并记录特征列顺序:
# 替换原LabelEncoder循环,为每个分类列单独创建并保存编码器 cols = ['DistributionChannel', 'MarketSegment'] # 处理DistributionChannel le_dist = LabelEncoder() df['DistributionChannel'] = le_dist.fit_transform(df['DistributionChannel']) pickle.dump(le_dist, open('le_dist.pkl', 'wb')) # 处理MarketSegment le_market = LabelEncoder() df['MarketSegment'] = le_market.fit_transform(df['MarketSegment']) pickle.dump(le_market, open('le_market.pkl', 'wb')) # 保存数值缩放器 scaler = MinMaxScaler() X[cols_to_scale] = scaler.fit_transform(X[cols_to_scale]) pickle.dump(scaler, open('scaler.pkl', 'wb')) # 保存训练时的特征列顺序,确保部署时输入顺序一致 pickle.dump(X.columns.tolist(), open('feature_names.pkl', 'wb'))
步骤2:修改部署代码,正确应用预处理
更新app.py,加载所有预处理组件,并对输入特征做与训练时完全一致的转换:
import numpy as np from flask import Flask, request, jsonify, render_template import pickle app = Flask(__name__) # 加载模型和所有预处理组件 model = pickle.load(open('model1.pkl','rb')) le_dist = pickle.load(open('le_dist.pkl', 'rb')) le_market = pickle.load(open('le_market.pkl', 'rb')) scaler = pickle.load(open('scaler.pkl', 'rb')) feature_names = pickle.load(open('feature_names.pkl', 'rb')) cols_to_scale = ['Age','DaysSinceCreation', 'AverageLeadTime', 'LodgingRevenue', 'OtherRevenue','DaysSinceLastStay', 'DaysSinceFirstStay'] # 获取需要缩放的特征在列表中的索引 scale_indices = [feature_names.index(col) for col in cols_to_scale] @app.route('/') def home(): return render_template('index.html') @app.route('/predict',methods=['POST']) def predict(): # 获取表单输入,顺序必须与feature_names完全一致 features = [x for x in request.form.values()] # 转换分类特征 features[feature_names.index('DistributionChannel')] = le_dist.transform([features[feature_names.index('DistributionChannel')]])[0] features[feature_names.index('MarketSegment')] = le_market.transform([features[feature_names.index('MarketSegment')]])[0] # 转换为数组 processed_features = np.array(features, dtype=float) # 缩放数值特征 processed_features[scale_indices] = scaler.transform([processed_features[scale_indices]])[0] # 转为模型要求的2D数组格式 final_features = [processed_features] prediction = model.predict(final_features) # 分类问题直接输出类别,无需round output = "Yes" if prediction[0] == 1 else "No" return render_template('index.html', prediction_text=f'Booked: {output}') if __name__ == "__main__": app.run(debug=True)
步骤3:确认网页表单的输入字段
检查index.html中的表单,确保输入字段的数量为26个,且顺序与feature_names完全一致,每个字段对应训练时X的一列。
验证
重新运行model.py生成新的模型和预处理文件,启动Flask服务后测试预测,即可解决特征数不匹配的问题。
内容的提问来源于stack exchange,提问作者Karthik Bhandary
相关产品推荐
相关产品推荐

