为何无法使用pickle序列化自定义__dict__属性的dataclass?
__dict__ property using asdict I defined a simple dataclass with a custom __dict__ property that uses asdict, but pickle can't serialize it. Here's my code:
import pickle from dataclasses import dataclass, asdict @dataclass class Point: x: int y: int @property def __dict__(self): return asdict(self) p = Point(10, 20) assert p.__dict__ == {'x': 10, 'y': 20} print(p.__dict__) b = pickle.dumps(p) p2 = pickle.loads(b) print(p2)
When I run this, I get the following output:
{'x': 10, 'y': 20} Traceback (most recent call last): File "C:\Projects\others\pythonPlayground\data_class.py", line 18, in <module> p2 = pickle.loads(b) File "C:\Projects\others\pythonPlayground\data_class.py", line 11, in __dict__ return asdict(self) File "C:\ProgramData\Anaconda3\envs\pythonPlayground\lib\dataclasses.py", line 1075, in asdict return _asdict_inner(obj, dict_factory) File "C:\ProgramData\Anaconda3\envs\pythonPlayground\lib\dataclasses.py", line 1082, in _asdict_inner value = _asdict_inner(getattr(obj, f.name), dict_factory) AttributeError: 'Point' object has no attribute 'x'
Why is this happening? How can I fix it? Are there other general-purpose serialization libraries that can byte-serialize almost any object? I tried dill and cloudpickle—dill works, but cloudpickle doesn't.
Let's break down what's going on here and how to fix it.
Root Cause
The problem stems from how pickle deserializes objects and how your custom __dict__ property clashes with that workflow:
- When pickle loads an object, it first creates an empty instance without running
__init__or setting any attributes. - Next, it tries to restore the object's state by populating its
__dict__. But since you've overridden__dict__as a property that callsasdict(self),asdicttries to accessxandy—which haven't been initialized yet in the partially built instance. That's exactly why you get theAttributeError. - On top of that, dataclasses already manage their own
__dict__internally, so overriding it as a property breaks the normal serialization/deserialization cycle pickle relies on.
Fixes
Here are a few clean, practical solutions:
1. Remove the custom __dict__ property (recommended)
You don't need to override __dict__ to get a dictionary representation of your dataclass—just call asdict(p) directly when you need it. The default dataclass implementation handles __dict__ correctly for pickle:
import pickle from dataclasses import dataclass, asdict @dataclass class Point: x: int y: int p = Point(10, 20) # Use asdict() explicitly when you need the dict format assert asdict(p) == {'x': 10, 'y': 20} print(asdict(p)) b = pickle.dumps(p) p2 = pickle.loads(b) print(p2) # Output: Point(x=10, y=20)
2. Rename the property (if you need a dict accessor)
If you really want a property that returns the dataclass as a dict, name it something other than __dict__—like to_dict—to avoid conflicting with Python's internal attribute handling:
import pickle from dataclasses import dataclass, asdict @dataclass class Point: x: int y: int @property def to_dict(self): return asdict(self) p = Point(10, 20) assert p.to_dict == {'x': 10, 'y': 20} print(p.to_dict) b = pickle.dumps(p) p2 = pickle.loads(b) print(p2) # Output: Point(x=10, y=20) print(p2.to_dict) # Output: {'x': 10, 'y': 20}
3. Implement custom pickle methods (advanced)
If you must keep the __dict__ property, define __getstate__ and __setstate__ to take control of how pickle handles your object's state:
import pickle from dataclasses import dataclass, asdict @dataclass class Point: x: int y: int @property def __dict__(self): return asdict(self) def __getstate__(self): # Return the data needed to rebuild the object return asdict(self) def __setstate__(self, state): # Manually set attributes from the saved state self.x = state['x'] self.y = state['y'] p = Point(10, 20) assert p.__dict__ == {'x': 10, 'y': 20} print(p.__dict__) b = pickle.dumps(p) p2 = pickle.loads(b) print(p2) # Output: Point(x=10, y=20)
Alternative Serialization Libraries
You mentioned testing dill and cloudpickle—here's why they behave differently, plus other options:
- dill: Works because it's built to handle more edge cases than standard pickle, including objects with custom
__dict__properties. It's a great drop-in replacement for most pickle use cases. - cloudpickle: Fails here because it still relies on core pickle mechanisms that your custom
__dict__disrupts. It's better suited for serializing things like lambda functions or objects from interactive sessions rather than this specific edge case.
Other worth-considering libraries:
- msgpack: Fast binary serialization that pairs well with dataclasses (convert to a dict with
asdictfirst, then serialize). - pydantic: If you're open to using a dataclass alternative, Pydantic models have built-in serialization (
.dict()and.json()) and play nicely with pickle and other tools.
内容的提问来源于stack exchange,提问作者Matěj Račinský

