Python中pickle与dill类定义更新行为差异及原因问询
为什么dill未沿用pickle的类更新行为?
dill在反序列化时只会更新被序列化对象自身的类定义,不会更新其包含的嵌套对象的类;而pickle会同时更新顶层对象和嵌套对象的类定义。两者的行为差异根源在于序列化时对类引用的处理逻辑不同。
pickle实验
实验代码
import os import pickle import tempfile from dataclasses import dataclass, field def pickle_save(x): with tempfile.NamedTemporaryFile(delete=False) as f: pickle.dump(x, f) return f def pickle_load(f): with open(f.name, "rb") as f: x = pickle.load(f) os.unlink(f.name) return x @dataclass class B: attribute: str = "old" def method_1(self): print(f"old class: {self.attribute=}") @dataclass class A: attribute_1: str = "old" instances_of_B: list[B] = field(default_factory=list) def method_1(self): print(f"old class: {self.attribute_1=}, {self.instances_of_B=}") def add_b_instance(self): self.instances_of_B.append(B()) old_a = A() old_a.add_b_instance() old_a.method_1() old_a.instances_of_B[0].method_1() print(f"{old_a = }") temp_file = pickle_save(old_a) # 已将old_a保存到文件,接下来更新类定义 # 再从文件加载old_a,验证新增方法是否可用 @dataclass class A: attribute_1: str = "new" attribute_2: str = "new attribute 2" instances_of_B: list[B] = field(default_factory=list) def method_1(self): print(f"new class: {self.attribute_1=}, {self.instances_of_B=}") def method_2(self): print("this method from A did not exist before") print(f"this attribute did not exist before: {self.attribute_2=}") @dataclass class B: attribute: str = "new" def method_1(self): print(f"new class: {self.attribute=}") def method_2(self): print("this method from B did not exist before") new_a = pickle_load(temp_file) print(f"{new_a=}") new_a.method_1() new_a.method_2() new_a.instances_of_B[0].method_1() new_a.instances_of_B[0].method_2()
输出结果
old class: self.attribute_1='old', self.instances_of_B=[B(attribute='old')] old class: self.attribute='old' old_a = A(attribute_1='old', instances_of_B=[B(attribute='old')]) new_a=A(attribute_1='old', attribute_2='new attribute 2', instances_of_B=[B(attribute='old')]) new class: self.attribute_1='old', self.instances_of_B=[B(attribute='old')] this method from A did not exist before this attribute did not exist before: self.attribute_2='new attribute 2' new class: self.attribute='old' this method from B did not exist before
从结果可见,反序列化后的A实例和其包含的B实例都能调用新增的method_2,说明pickle在加载时,会将顶层对象和嵌套对象都关联到当前环境中更新后的类定义。
dill实验
实验代码
只需替换pickle的导入语句,其余代码与pickle实验完全一致:
import dill as pickle
输出结果
old class: self.attribute_1='old', self.instances_of_B=[B(attribute='old')] old class: self.attribute='old' old_a = A(attribute_1='old', instances_of_B=[B(attribute='old')]) new_a=A(attribute_1='old', attribute_2='new attribute 2', instances_of_B=[B(attribute='old')]) new class: self.attribute_1='old', self.instances_of_B=[B(attribute='old')] this method from A did not exist before this attribute did not exist before: self.attribute_2='new attribute 2' old class: self.attribute='old' Traceback (most recent call last): File "c:\question_dill_pickle.py", line 78, in <module> new_a.instances_of_B[0].method_2() AttributeError: 'B' object has no attribute 'method_2'
结果显示,只有顶层A实例能使用新增方法,嵌套的B实例调用method_2时触发AttributeError,说明dill仅更新了顶层对象的类定义,嵌套对象仍关联旧的类定义。
行为差异的核心原因
pickle和dill的处理逻辑差异源于设计目标的不同:
- pickle默认使用全局名称引用:序列化对象时,只记录类的全局名称(比如
__main__.B),反序列化时直接从当前环境中查找该名称对应的类。因此无论顶层还是嵌套对象,都会自动关联到当前环境中最新的类定义。 - dill为了支持更复杂的序列化场景(比如序列化未在全局命名空间定义的类、闭包类、动态生成的类等),默认会序列化类的完整定义或内部引用,而非仅依赖全局名称。对于嵌套对象,dill在序列化时会将B类的旧定义和A实例绑定在一起,反序列化时优先使用这个内嵌的旧类定义,而非当前环境中更新后的B类。
简言之,pickle依赖全局命名空间实现类的自动更新,而dill为了兼容性和扩展性,选择了更保守的类绑定策略,避免因全局环境变化导致反序列化失败,代价是嵌套对象无法自动同步到新类。
内容的提问来源于stack exchange,提问作者thomast
相关产品推荐
相关产品推荐

