You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flask本地电动车价格预测应用提交表单后出现ValueError(未知ZIP编码类别)

Flask本地电动车价格预测应用提交表单后出现ValueError(未知ZIP编码类别)

项目背景

我正在做一个机器学习项目,目标是预测美国不同地区的电动车价格,主要是想巩固自己的实操能力。目前已经完成了独热编码、模型训练,还在本地跑通了Flask应用。

刚才我在本地表单里填了以下信息并提交:

County: Jefferson
City: PORT TOWNSEND
ZIP Code: 98368
Model Year: 2012
Make: NISSAN
Model: LEAF
Electric Vehicle Type: Battery Electric Vehicle (BEV)
CAFV Eligibility: Clean Alternative Fuel Vehicle Eligible
Legislative District: 24

遇到的问题

提交后直接弹出了ValueError错误,具体报错信息如下:

ValueError
ValueError: Found unknown categories ['98368'] in column 2 during transform

Traceback (most recent call last)
File "C:\Users\austin.conda\envs\electric_vehicle_price_prediction_2\lib\site-packages\flask\app.py", line 1498, in __call__
return self.wsgi_app(environ, start_response)
File "C:\Users\austin.conda\envs\electric_vehicle_price_prediction_2\lib\site-packages\flask\app.py", line 1476, in wsgi_app
response = self.handle_exception(e)
File "C:\Users\austin.conda\envs\electric_vehicle_price_prediction_2\lib\site-packages\flask\app.py", line 1473, in wsgi_app
response = self.full_dispatch_request()
File "C:\Users\austin.conda\envs\electric_vehicle_price_prediction_2\lib\site-packages\flask\app.py", line 882, in full_dispatch_request
rv = self.handle_user_exception(e)
File "C:\Users\austin.conda\envs\electric_vehicle_price_prediction_2\lib\site-packages\flask\app.py", line 880, in full_dispatch_request
rv = self.dispatch_request()
File "C:\Users\austin.conda\envs\electric_vehicle_price_prediction_2\lib\site-packages\flask\app.py", line 865, in dispatch_request
return self.ensure_sync(self.view_functions[rule.endpoint])(**view_args)  # type: ignore[no-any-return]
File "G:\Machine_Learning_Projects\austin\electric_vehicle_price_prediction_2\app\routes.py", line 38, in predict
price = predict_price(features)
File "G:\Machine_Learning_Projects\austin\electric_vehicle_price_prediction_2\app\model.py", line 29, in predict_price
transformed_features = encoder.transform(features_df)
File "C:\Users\austin.conda\envs\electric_vehicle_price_prediction_2\lib\site-packages\sklearn\utils_set_output.py", line 157, in wrapped
data_to_wrap = f(self, X, *args, **kwargs)
File "C:\Users\austin.conda\envs\electric_vehicle_price_prediction_2\lib\site-packages\sklearn\preprocessing_encoders.py", line 1027, in transform
X_int, X_mask = self._transform(
File "C:\Users\austin.conda\envs\electric_vehicle_price_prediction_2\lib\site-packages\sklearn\preprocessing_encoders.py", line 200, in _transform
raise ValueError(msg)
ValueError: Found unknown categories ['98368'] in column 2 during transform

我尝试过的代码

app/routes.py 文件代码

from flask import render_template, request, jsonify
from app import app
from app.model import predict_price
from jinja2 import Environment, FileSystemLoader, PackageLoader, select_autoescape

@app.route('/')
def index():
env = Environment(
loader=PackageLoader("app"),
autoescape=select_autoescape()
)
template = env.get_template("index.html")
return render_template(template)

@app.route('/predict', methods=['POST'])
def predict():
data = request.form.to_dict()

    # Convert the form data into the correct format for prediction
    features = [
        data['county'],
        data['city'],
        data['zip_code'],
        data['model_year'],
        data['make'],
        data['model'],
        data['ev_type'],
        data['cafv_eligibility'],
        data['legislative_district']
    ]
    
    # Get the prediction result
    price = predict_price(features)
    
    return jsonify({'predicted_price': price})

app/model.py 文件代码

import pandas as pd
from sklearn.preprocessing import OneHotEncoder
from sklearn.ensemble import RandomForestRegressor
import joblib
from flask import Flask, render_template
from jinja2 import Environment, FileSystemLoader, PackageLoader, select_autoescape

env = Environment(
loader=PackageLoader("app"),
autoescape=select_autoescape()
)

model = joblib.load('model/ev_price_model.pkl')

def predict_price(features):
    encoder = joblib.load('model/encoder.pkl')  # Load encoder if needed
    
    features_df = pd.DataFrame([features], columns=['County', 'City', 'ZIP Code', 'Model Year', 'Make', 'Model', 'Electric Vehicle Type', 'Clean Alternative Fuel Vehicle (CAFV) Eligibility', 'Legislative District'])
    
    # Apply encoding, scaling, etc., if necessary
    transformed_features = encoder.transform(features_df)
    
    # Make the prediction
    price = model.predict(transformed_features)
    
    return price[0]  # Assuming it returns a single value

我的期望

我本来以为已经完成了独热编码,提交表单后应该能正常拿到预测结果,没想到会出这个问题,希望有人能帮我解决。


解决方案

这个问题的根源很明确:你训练独热编码器(encoder.pkl)的时候,训练数据集里压根没出现过98368这个ZIP编码,所以预测时编码器碰到陌生类别直接报错了。下面给你几个实用的解决办法:

方法1:让编码器忽略未知类别

在训练编码器的时候,加上handle_unknown='ignore'参数,这样碰到训练时没见过的类别,编码器会自动忽略它(对应编码列全设为0),不会抛出错误。

训练阶段的代码要改成这样:

encoder = OneHotEncoder(handle_unknown='ignore')
# 用训练数据拟合编码器
encoder.fit(train_data[['County', 'City', 'ZIP Code', ...]])
# 重新保存编码器
joblib.dump(encoder, 'model/encoder.pkl')

注意:改完后需要重新训练并替换原来的encoder.pkl文件。

方法2:合并稀有ZIP编码类别

如果你的训练数据里有些ZIP编码出现次数极少(比如只出现一两次),可以把这些稀有类别统一合并成“Other”,这样既能减少编码维度,又能避免预测时碰到陌生类别。

预处理数据时可以这么做:

# 统计每个ZIP编码的出现次数
zip_counts = train_data['ZIP Code'].value_counts()
# 设置阈值,比如只保留出现次数≥5的ZIP编码
threshold = 5
# 把稀有ZIP编码替换成'Other'
train_data['ZIP Code'] = train_data['ZIP Code'].apply(lambda x: x if zip_counts[x] >= threshold else 'Other')

之后再重新训练编码器,预测时碰到陌生ZIP编码,就先把它替换成'Other'再传入编码器。

方法3:预测前主动检查并处理未知ZIP编码

在model.py的predict_price函数里,先判断输入的ZIP编码是否在编码器的已知类别里,如果不在,就替换成一个常见类别或者“Other”。

修改后的predict_price函数示例:

def predict_price(features):
    encoder = joblib.load('model/encoder.pkl')
    
    features_df = pd.DataFrame([features], columns=['County', 'City', 'ZIP Code', 'Model Year', 'Make', 'Model', 'Electric Vehicle Type', 'Clean Alternative Fuel Vehicle (CAFV) Eligibility', 'Legislative District'])
    
    # 获取编码器中ZIP Code列的所有已知类别(ZIP是第3列,索引从0开始)
    zip_categories = encoder.categories_[2]
    input_zip = features_df['ZIP Code'].iloc[0]
    
    # 如果输入的ZIP不在已知类别里,替换成训练数据中出现最多的ZIP编码
    if input_zip not in zip_categories:
        # 这里需要你提前知道训练数据里最常见的ZIP,或者替换成'Other'
        features_df['ZIP Code'] = 'Other'
    
    transformed_features = encoder.transform(features_df)
    price = model.predict(transformed_features)
    
    return price[0]

额外提醒

还要确认训练时ZIP编码的数据类型和预测时输入的类型一致,比如训练时是字符串,预测时输入的也必须是字符串,别因为类型不一致导致误判成“未知类别”。


备注:内容来源于stack exchange,提问作者Steve Austin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 16:29:31