Python如何实现类似Boost的格式无关类序列化?
Great question! Coming from C++ Boost Serialization, it's totally reasonable to want that same format-agnostic, decentralized approach in Python—ditching the messy if-else chains is a must. Let's walk through a clean implementation that checks all your boxes, and even hits those optional ones.
Core Idea
We'll replicate Boost's pattern with two key components:
- Abstract Archive Base Class: Defines a common interface for all serialization formats (JSON, binary, XML, etc.).
- Registration System: Maps custom classes to their serialization logic, so we don't need centralized if-else checks.
Step 1: Build the Archive Foundation
First, create an abstract Archive class that handles dispatching to class-specific serialization functions. We'll also add a registry to map types to their handlers.
# Core archive infrastructure _serialization_registry = {} def register_serializer(obj_type): """Decorator to register a serialization function for a class.""" def decorator(func): _serialization_registry[obj_type] = func return func return decorator class Archive: def serialize(self, obj, version=0): """Dispatch to the registered serialization function for the object's type.""" serialize_func = _serialization_registry.get(type(obj)) if not serialize_func: raise TypeError(f"No serializer registered for {type(obj).__name__}") serialize_func(self, obj, version) def __and__(self, value): """Overload & operator to mimic Boost's clean syntax.""" raise NotImplementedError("Subclasses must implement __and__")
Step 2: Implement a Format-Specific Archive
Let's build a JSON archive first—you can replicate this pattern for binary, XML, or any other format later. This archive will handle both reading and writing using a mode flag, just like Boost's input/output archives.
import json class JSONArchive(Archive): def __init__(self, file_path, mode='w'): self.mode = mode self.file = open(file_path, mode, encoding='utf-8') if mode == 'r': self._data = json.load(self.file) self._index = 0 else: self._data = [] def __enter__(self): return self def __exit__(self, exc_type, exc_val, exc_tb): if self.mode == 'w': json.dump(self._data, self.file, indent=2) self.file.close() # Handle basic types (extend these for bool, float, list, etc.) def _handle_int(self, value): if self.mode == 'w': self._data.append(('int', value)) else: tag, val = self._data[self._index] assert tag == 'int', f"Expected int, got {tag}" self._index += 1 return val def _handle_str(self, value): if self.mode == 'w': self._data.append(('str', value)) else: tag, val = self._data[self._index] assert tag == 'str', f"Expected str, got {tag}" self._index += 1 return val def _handle_custom_type(self, value): if self.mode == 'r': # Create empty instance for deserialization obj = type(value).__new__(type(value)) self.serialize(obj) return obj else: self.serialize(value) return self def __and__(self, value): """Overload & to handle both reading and writing.""" if isinstance(value, int): return self._handle_int(value) elif isinstance(value, str): return self._handle_str(value) else: return self._handle_custom_type(value)
Step 3: Register Your Classes
Now, you can register serialization logic for your classes—no changes to the class itself (non-intrusive), and you can do this in the same file as the class (decentralized).
# Your custom classes (no modifications needed!) class Foo: def __init__(self, some_number=0, some_string=""): self.some_number = some_number self.some_string = some_string class Bar: def __init__(self, bar_value=0): self.bar_value = bar_value class NestedFoo: def __init__(self, foo=None, bar=None): self.foo = foo or Foo() self.bar = bar or Bar() # Register serialization handlers (decentralized—put these next to your classes) @register_serializer(Foo) def serialize_foo(ar, foo, version): # Use & operator, just like Boost! foo.some_number = ar & foo.some_number foo.some_string = ar & foo.some_string @register_serializer(Bar) def serialize_bar(ar, bar, version): bar.bar_value = ar & bar.bar_value @register_serializer(NestedFoo) def serialize_nested_foo(ar, nested_foo, version): # Nested types work automatically—no extra code needed nested_foo.foo = ar & nested_foo.foo nested_foo.bar = ar & nested_foo.bar
Step 4: Test It Out
Use the archive just like you would in Boost—switch formats by swapping the archive class, not your serialization logic.
# Serialize to JSON with JSONArchive("test.json", "w") as ar: my_foo = Foo(42, "Hello Boost!") my_nested = NestedFoo(my_foo, Bar(100)) ar & my_nested # Deserialize from JSON with JSONArchive("test.json", "r") as ar: loaded_nested = NestedFoo() loaded_nested = ar & loaded_nested print(loaded_nested.foo.some_number) # Output: 42 print(loaded_nested.bar.bar_value) # Output: 100
Does This Meet Your Requirements?
Let's verify:
- Decentralized: Add new classes by registering their serializer in their own file—no central if-else mess.
- Nested Classes: Works out of the box—serializing
NestedFooautomatically usesFooandBar's registered handlers. - Read/Write Same Method: The
&operator handles both directions, just like Boost's archives. - Non-Intrusive: No changes to your original
Foo,Bar, orNestedFooclasses—all logic is in separate serializer functions.
Bonus: Extend to Other Formats
To add a binary archive, just create a new subclass of Archive that implements __and__ with binary handling logic. Your existing class serializers will work without any changes—true format agnosticism!
内容的提问来源于stack exchange,提问作者pschill

