Python 3.6多进程更新含不可序列化对象字典的性能优化问询
Let's break down how to fix this serialization overhead killing your multiprocessing gains. I’ve dealt with similar headaches with unpickleable objects before, so here are some practical, actionable solutions tailored to your Python 3.6 setup:
1. Use Shared Memory Structures to Avoid Full List Serialization
The biggest hit to performance is likely serializing/deserializing entire lists to pass them between processes. Instead, use process-safe shared objects so child processes can modify the dictionary and lists directly without copying or re-serializing everything.
Since you’re using pathos.multiprocessing, you can still leverage Python’s built-in multiprocessing.Manager to create shared dicts and lists:
from pathos.multiprocessing import ProcessPool from multiprocessing import Manager def add_items_to_shared_dict(shared_dict, key, item_params): # Generate your unpickleable objects directly in the child process (even better!) new_items = generate_your_unpickleable_objects(item_params) # Modify the shared list directly—no need to send the whole list back shared_dict[key].extend(new_items) if __name__ == "__main__": # Create a manager for shared objects manager = Manager() shared_dict = manager.dict() # Initialize empty shared lists for each key you need shared_dict["dataset1"] = manager.list() shared_dict["dataset2"] = manager.list() # Set up your pool pool = ProcessPool() # Submit tasks—only send params, not the unpickleable objects themselves tasks = [ (shared_dict, "dataset1", {"param_a": 10, "param_b": 20}), (shared_dict, "dataset2", {"param_a": 30, "param_b": 40}) ] pool.map(lambda args: add_items_to_shared_dict(*args), tasks) # Convert shared lists back to regular lists if needed final_dict = {k: list(v) for k, v in shared_dict.items()}
Why this works: The manager handles inter-process communication under the hood, and you only serialize small parameter sets instead of entire lists of unpickleable objects.
2. Generate Unpickleable Objects Inside Child Processes
If you’re currently creating objects in the main process and trying to pass them to child processes, that’s a double whammy: you have to serialize them to send over, then re-serialize to add to the dict. Instead, shift the object generation logic into the child processes entirely.
This way, you never have to serialize the unpickleable objects at all—you just pass the parameters needed to create them, and the child process builds and adds them directly to the shared structure (as shown in the example above). This is the single most impactful fix for this scenario.
3. Use Fork-Based Multiprocessing (Unix Only)
If you’re running on Linux/macOS (Unix systems), Python 3.6 defaults to fork for spawning child processes. Forking copies the main process’s memory space (using copy-on-write), so child processes can access the main process’s dictionary directly—no serialization required. Just add a lock to ensure process safety:
from pathos.multiprocessing import ProcessPool import multiprocessing as mp # Define your shared dict and lock in the main process shared_dict = {"dataset1": []} lock = mp.Lock() def add_items(key, item_params): new_items = generate_your_unpickleable_objects(item_params) # Use a lock to prevent race conditions when modifying the shared dict with lock: shared_dict[key].extend(new_items) if __name__ == "__main__": # Force fork context (though it's default on Unix) pool = ProcessPool(context=mp.get_context("fork")) pool.apply_async(add_items, args=("dataset1", {"param_a": 10})) pool.close() pool.join()
Note: This won’t work on Windows, which uses spawn instead of fork (spawn doesn’t inherit the main process’s memory).
4. Optimize Serialization for Unpickleable Objects
If you absolutely have to pass objects between processes, try these tweaks to reduce serialization overhead:
- Use
cloudpickleinstead ofdill(it’s optimized for more edge cases and can be faster for certain object types). - Implement a custom
__reduce__method for your unpickleable classes. This lets you control exactly what gets serialized (e.g., skip non-serializable attributes and recreate them in the child process):import dill class UnpickleableObject: def __init__(self, data, non_serializable_resource): self.data = data self.resource = non_serializable_resource def __reduce__(self): # Serialize only the data, not the resource return (UnpickleableObject, (self.data, None), {"resource": self._recreate_resource()}) def _recreate_resource(self): # Logic to rebuild the non-serializable resource in the child process return create_your_resource()
This cuts down on the amount of data that needs to be serialized, speeding up inter-process communication.
内容的提问来源于stack exchange,提问作者fortea

