You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Streamlit加载Pickle序列化Airbnb价格预测模型报错求助

新加坡Airbnb价格预测模型Streamlit加载时Pickle反序列化报错

问题背景

我构建了用于新加坡Airbnb价格预测的机器学习模型,使用Pickle在Python 3.9中序列化训练好的模型并推送至GitHub。对应的Streamlit应用曾正常运行,近期加载模型时出现如下报错:

Traceback (most recent call last):
      File "/home/appuser/venv/lib/python3.9/site-packages/streamlit/runtime/scriptrunner/script_runner.py", line 552, in _run_script
        exec(code, module.__dict__)
      File "/app/singapore-airbnb/new.py", line 119, in <module>
        rf = pickle.load(file)
      File "sklearn/tree/_tree.pyx", line 714, in sklearn.tree._tree.Tree.__setstate__
      File "sklearn/tree/_tree.pyx", line 1418, in sklearn.tree._tree._check_node_ndarray
    ValueError: node array from the pickle has an incompatible dtype:
    - expected: {'names': ['left_child', 'right_child', 'feature', 'threshold', 'impurity', 'n_node_samples', 'weighted_n_node_samples', 'missing_go_to_left'], 'formats': ['<i8', '<i8', '<i8', '<f8', '<f8', '<i8', '<f8', 'u1'], 'offsets': [0, 8, 16, 24, 32, 40, 48, 56], 'itemsize': 64}
    - got     : [('left_child', '<i8'), ('right_child', '<i8'), ('feature', '<i8'), ('threshold', '<f8'), ('impurity', '<f8'), ('n_node_samples', '<i8'), ('weighted_n_node_samples', '<f8')]

/home/appuser/venv/lib/python3.9/site-packages/sklearn/base.py:347: InconsistentVersionWarning: Trying to unpickle estimator DecisionTreeRegressor from version 1.2.0 when using version 1.3.0. This might lead to breaking code or invalid results. Use at your own risk. For more info please refer to:
https://scikit-learn.org/stable/model_persistence.html#security-maintainability-limitations
      warnings.warn(

问题原因

训练模型时使用的scikit-learn版本为1.2.0,而Streamlit部署环境当前使用的是1.3.0版本。scikit-learn 1.3.0对DecisionTreeRegressor的节点数组结构做了修改(新增missing_go_to_left字段),导致旧版本序列化的模型无法在新版本中正常反序列化。

解决方案

方案1:统一scikit-learn版本

  • 本地安装1.3.0版本的scikit-learn,重新训练模型并序列化后推送至GitHub,确保部署环境和训练环境版本一致
  • 或者在Streamlit项目的requirements.txt中添加scikit-learn==1.2.0,强制部署环境使用和训练时相同的版本

方案2:改用scikit-learn官方推荐的joblib序列化

scikit-learn官方建议用joblib替代pickle序列化树模型这类含大量数值数据的模型:

  • 训练保存模型时:
    import joblib
    joblib.dump(rf, 'singapore_airbnb_model.joblib')
    
  • 加载模型时:
    import joblib
    rf = joblib.load('singapore_airbnb_model.joblib')
    
    同样需要保证训练和部署环境的scikit-learn版本一致

内容的提问来源于stack exchange,提问作者George Ng

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 04:09:51