跨环境加载含时间戳的Pickle数据集报错:_unpickle_timestamp参数不匹配
问题:跨Python版本加载含时间戳的pickle Pandas DataFrame失败
问题背景
在Python 3.10.12环境(环境1)中构建并通过pickle存储了包含时间戳的Pandas DataFrame数据集,原环境可正常加载,但在Python 3.9.12环境(环境2)中加载时触发时间戳反序列化错误。因Keras依赖安装难度高,不愿升级环境2的Python版本,寻求可行解决方案。
报错信息
runfile('D:/driver_backprop.py', wdir='D:') Traceback (most recent call last): File "C:\Users\jerem\anaconda3\envs\keras\lib\site-packages\spyder_kernels\py3compat.py", line 356, in compat_exec exec(code, globals, locals) File "d:\driver_backprop.py", line 14, in <module> data_df = pickle.load(f) File "pandas\_libs\tslibs\timestamps.pyx", line 132, in pandas._libs.tslibs.timestamps._unpickle_timestamp TypeError: _unpickle_timestamp() takes exactly 3 positional arguments (4 given)
加载代码
import pickle with open('data_w-solar.pickle','rb') as f: data_df = pickle.load(f) station_dict = pickle.load(f)
环境版本信息
环境1(Python 3.10.12,可正常加载)
(shadow) C:\Users\jerem>conda list pickle # packages in environment at C:\Users\jerem\anaconda3\envs\shadow: # # Name Version Build Channel cloudpickle 2.2.1 py310haa95532_0 pickleshare 0.7.5 pyhd3eb1b0_1003
环境2(Python 3.9.12,加载失败)
(keras) C:\Users\jerem>conda list pickle # packages in environment at C:\Users\jerem\anaconda3\envs\keras: # # Name Version Build Channel cloudpickle 2.2.1 py39haa95532_0 pickleshare 0.7.5 pyhd3eb1b0_1003
两个环境的pickle相关包版本一致,但cloudpickle的Build版本不同,尝试安装环境1对应的cloudpickle Build版本但Conda无法找到解决方案。
解决方案
方案1:改用Pandas内置序列化方法(推荐)
在环境1中重新保存数据时,使用Pandas针对自身结构优化的序列化方法,替代原生pickle,从根源避免跨版本时间戳兼容性问题:
import pandas as pd import pickle # 重新保存DataFrame和字典 data_df.to_pickle('data_w-solar_pandas.pkl') with open('station_dict.pkl', 'wb') as f: pickle.dump(station_dict, f) # 环境2中加载 data_df = pd.read_pickle('data_w-solar_pandas.pkl') with open('station_dict.pkl', 'rb') as f: station_dict = pickle.load(f)
方案2:自定义时间戳还原函数
在环境2中加载前,注册自定义的时间戳还原逻辑,适配参数数量差异:
import pickle from pandas._libs.tslibs.timestamps import Timestamp def custom_unpickle_timestamp(year, month, day, nanos=None): # 兼容4个参数的传入情况,忽略多余参数 return Timestamp(year=year, month=month, day=day, nanos=nanos if nanos is not None else 0) # 替换pickle的默认还原函数 pickle.Unpickler.dispatch[Timestamp] = custom_unpickle_timestamp # 正常加载数据 with open('data_w-solar.pickle','rb') as f: data_df = pickle.load(f) station_dict = pickle.load(f)
方案3:统一Pandas版本
检查两个环境的Pandas版本是否一致,若环境2的Pandas版本过低,升级到与环境1相同的版本(无需升级Python):
# 先查看环境1的Pandas版本 conda list pandas # 在环境2中安装对应版本 conda install pandas=X.X.X
Pandas对时间戳的序列化逻辑在不同版本中有调整,统一版本可解决大部分跨环境序列化问题。
内容的提问来源于stack exchange,提问作者Jeremy Matt
相关产品推荐
相关产品推荐

