You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3.6多进程更新含不可序列化对象字典的性能优化问询

Let's break down how to fix this serialization overhead killing your multiprocessing gains. I’ve dealt with similar headaches with unpickleable objects before, so here are some practical, actionable solutions tailored to your Python 3.6 setup:

1. Use Shared Memory Structures to Avoid Full List Serialization

The biggest hit to performance is likely serializing/deserializing entire lists to pass them between processes. Instead, use process-safe shared objects so child processes can modify the dictionary and lists directly without copying or re-serializing everything.

Since you’re using pathos.multiprocessing, you can still leverage Python’s built-in multiprocessing.Manager to create shared dicts and lists:

from pathos.multiprocessing import ProcessPool
from multiprocessing import Manager

def add_items_to_shared_dict(shared_dict, key, item_params):
    # Generate your unpickleable objects directly in the child process (even better!)
    new_items = generate_your_unpickleable_objects(item_params)
    # Modify the shared list directly—no need to send the whole list back
    shared_dict[key].extend(new_items)

if __name__ == "__main__":
    # Create a manager for shared objects
    manager = Manager()
    shared_dict = manager.dict()
    
    # Initialize empty shared lists for each key you need
    shared_dict["dataset1"] = manager.list()
    shared_dict["dataset2"] = manager.list()

    # Set up your pool
    pool = ProcessPool()
    
    # Submit tasks—only send params, not the unpickleable objects themselves
    tasks = [
        (shared_dict, "dataset1", {"param_a": 10, "param_b": 20}),
        (shared_dict, "dataset2", {"param_a": 30, "param_b": 40})
    ]
    pool.map(lambda args: add_items_to_shared_dict(*args), tasks)

    # Convert shared lists back to regular lists if needed
    final_dict = {k: list(v) for k, v in shared_dict.items()}

Why this works: The manager handles inter-process communication under the hood, and you only serialize small parameter sets instead of entire lists of unpickleable objects.

2. Generate Unpickleable Objects Inside Child Processes

If you’re currently creating objects in the main process and trying to pass them to child processes, that’s a double whammy: you have to serialize them to send over, then re-serialize to add to the dict. Instead, shift the object generation logic into the child processes entirely.

This way, you never have to serialize the unpickleable objects at all—you just pass the parameters needed to create them, and the child process builds and adds them directly to the shared structure (as shown in the example above). This is the single most impactful fix for this scenario.

3. Use Fork-Based Multiprocessing (Unix Only)

If you’re running on Linux/macOS (Unix systems), Python 3.6 defaults to fork for spawning child processes. Forking copies the main process’s memory space (using copy-on-write), so child processes can access the main process’s dictionary directly—no serialization required. Just add a lock to ensure process safety:

from pathos.multiprocessing import ProcessPool
import multiprocessing as mp

# Define your shared dict and lock in the main process
shared_dict = {"dataset1": []}
lock = mp.Lock()

def add_items(key, item_params):
    new_items = generate_your_unpickleable_objects(item_params)
    # Use a lock to prevent race conditions when modifying the shared dict
    with lock:
        shared_dict[key].extend(new_items)

if __name__ == "__main__":
    # Force fork context (though it's default on Unix)
    pool = ProcessPool(context=mp.get_context("fork"))
    pool.apply_async(add_items, args=("dataset1", {"param_a": 10}))
    
    pool.close()
    pool.join()

Note: This won’t work on Windows, which uses spawn instead of fork (spawn doesn’t inherit the main process’s memory).

4. Optimize Serialization for Unpickleable Objects

If you absolutely have to pass objects between processes, try these tweaks to reduce serialization overhead:

  • Use cloudpickle instead of dill (it’s optimized for more edge cases and can be faster for certain object types).
  • Implement a custom __reduce__ method for your unpickleable classes. This lets you control exactly what gets serialized (e.g., skip non-serializable attributes and recreate them in the child process):
    import dill
    
    class UnpickleableObject:
        def __init__(self, data, non_serializable_resource):
            self.data = data
            self.resource = non_serializable_resource
    
        def __reduce__(self):
            # Serialize only the data, not the resource
            return (UnpickleableObject, (self.data, None), {"resource": self._recreate_resource()})
    
        def _recreate_resource(self):
            # Logic to rebuild the non-serializable resource in the child process
            return create_your_resource()
    

This cuts down on the amount of data that needs to be serialized, speeding up inter-process communication.


内容的提问来源于stack exchange,提问作者fortea

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:33:22