You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中pickle与dill类定义更新行为差异及原因问询

为什么dill未沿用pickle的类更新行为?

dill在反序列化时只会更新被序列化对象自身的类定义,不会更新其包含的嵌套对象的类;而pickle会同时更新顶层对象和嵌套对象的类定义。两者的行为差异根源在于序列化时对类引用的处理逻辑不同。

pickle实验

实验代码

import os
import pickle
import tempfile
from dataclasses import dataclass, field


def pickle_save(x):
    with tempfile.NamedTemporaryFile(delete=False) as f:
        pickle.dump(x, f)
    return f


def pickle_load(f):
    with open(f.name, "rb") as f:
        x = pickle.load(f)
    os.unlink(f.name)
    return x


@dataclass
class B:
    attribute: str = "old"

    def method_1(self):
        print(f"old class: {self.attribute=}")


@dataclass
class A:
    attribute_1: str = "old"
    instances_of_B: list[B] = field(default_factory=list)

    def method_1(self):
        print(f"old class: {self.attribute_1=}, {self.instances_of_B=}")

    def add_b_instance(self):
        self.instances_of_B.append(B())


old_a = A()
old_a.add_b_instance()
old_a.method_1()
old_a.instances_of_B[0].method_1()
print(f"{old_a = }")
temp_file = pickle_save(old_a)

# 已将old_a保存到文件,接下来更新类定义
# 再从文件加载old_a,验证新增方法是否可用

@dataclass
class A:
    attribute_1: str = "new"
    attribute_2: str = "new attribute 2"
    instances_of_B: list[B] = field(default_factory=list)

    def method_1(self):
        print(f"new class: {self.attribute_1=}, {self.instances_of_B=}")

    def method_2(self):
        print("this method from A did not exist before")
        print(f"this attribute did not exist before: {self.attribute_2=}")


@dataclass
class B:
    attribute: str = "new"

    def method_1(self):
        print(f"new class: {self.attribute=}")

    def method_2(self):
        print("this method from B did not exist before")


new_a = pickle_load(temp_file)
print(f"{new_a=}")
new_a.method_1()
new_a.method_2()
new_a.instances_of_B[0].method_1()
new_a.instances_of_B[0].method_2()

输出结果

old class: self.attribute_1='old', self.instances_of_B=[B(attribute='old')]
old class: self.attribute='old'
old_a = A(attribute_1='old', instances_of_B=[B(attribute='old')])
new_a=A(attribute_1='old', attribute_2='new attribute 2', instances_of_B=[B(attribute='old')])
new class: self.attribute_1='old', self.instances_of_B=[B(attribute='old')]
this method from A did not exist before
this attribute did not exist before: self.attribute_2='new attribute 2'
new class: self.attribute='old'
this method from B did not exist before

从结果可见,反序列化后的A实例和其包含的B实例都能调用新增的method_2,说明pickle在加载时,会将顶层对象和嵌套对象都关联到当前环境中更新后的类定义。

dill实验

实验代码

只需替换pickle的导入语句,其余代码与pickle实验完全一致:

import dill as pickle

输出结果

old class: self.attribute_1='old', self.instances_of_B=[B(attribute='old')]
old class: self.attribute='old'
old_a = A(attribute_1='old', instances_of_B=[B(attribute='old')])
new_a=A(attribute_1='old', attribute_2='new attribute 2', instances_of_B=[B(attribute='old')])       
new class: self.attribute_1='old', self.instances_of_B=[B(attribute='old')]
this method from A did not exist before
this attribute did not exist before: self.attribute_2='new attribute 2'
old class: self.attribute='old'
Traceback (most recent call last):
  File "c:\question_dill_pickle.py", line 78, in <module>
    new_a.instances_of_B[0].method_2()
AttributeError: 'B' object has no attribute 'method_2'

结果显示,只有顶层A实例能使用新增方法,嵌套的B实例调用method_2时触发AttributeError,说明dill仅更新了顶层对象的类定义,嵌套对象仍关联旧的类定义。

行为差异的核心原因

pickle和dill的处理逻辑差异源于设计目标的不同:

  • pickle默认使用全局名称引用:序列化对象时,只记录类的全局名称(比如__main__.B),反序列化时直接从当前环境中查找该名称对应的类。因此无论顶层还是嵌套对象,都会自动关联到当前环境中最新的类定义。
  • dill为了支持更复杂的序列化场景(比如序列化未在全局命名空间定义的类、闭包类、动态生成的类等),默认会序列化类的完整定义或内部引用,而非仅依赖全局名称。对于嵌套对象,dill在序列化时会将B类的旧定义和A实例绑定在一起,反序列化时优先使用这个内嵌的旧类定义,而非当前环境中更新后的B类。

简言之,pickle依赖全局命名空间实现类的自动更新,而dill为了兼容性和扩展性,选择了更保守的类绑定策略,避免因全局环境变化导致反序列化失败,代价是嵌套对象无法自动同步到新类。


内容的提问来源于stack exchange,提问作者thomast

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 07:10:31