向子进程传递任意对象:SharedNumpyArray类实例传递行为探究
Great question! Let's break down exactly what happens when you pass an instance of your SharedNumpyArray class to a child process, step by step, and highlight key behaviors you need to know.
Core Background: How Python Passes Objects Between Processes
By default, Python's multiprocessing module uses pickle serialization to send objects between processes. This means it converts the object into a byte stream, sends that to the child, then reconstructs the object from the byte stream in the child's address space.
Step-by-Step Behavior for SharedNumpyArray
1. Serialization (Parent Process)
When you pass your SharedNumpyArray instance to a child process, pickle will serialize all its public and private attributes that are pickle-compatible:
- Basic metadata:
shape,dtype,buf_sizeare simple values that serialize without issue. - Critical shared memory handle:
self.__buf(aRawArrayfrommultiprocessing.sharedctypes) is designed to be pickle-safe. Instead of copying the entire buffer's contents, pickle serializes a reference to the underlying shared memory segment.
2. Deserialization & Memory Mapping (Child Process)
Once the child process receives the serialized data:
- It reconstructs the
SharedNumpyArrayinstance, restoring the metadata and the__bufreference to the shared memory. - The
init_buf()method runs again, callingnp.frombuffer(self.__buf, dtype=self.dtype).reshape(self.shape). This creates a new numpy array view that maps directly to the same shared memory segment as the parent'sbuf.
3. Key Outcome: True Shared Memory
The child's buf and the parent's buf are views into the exact same block of shared memory. This means:
- Any changes made to
bufin the child process are immediately visible in the parent process (and vice versa). - No full copy of the array data is made between processes—this is extremely efficient for large arrays.
Critical Notes & Gotchas
- Avoid accidental copies: If you call
arr.buf.copy()in the child, you'll create a private copy of the data in the child's memory. Changes to this copy won't propagate back to the shared memory. - Process safety isn't automatic: Numpy arrays don't have built-in process synchronization. If multiple processes are reading/writing to
bufsimultaneously, you'll get race conditions (corrupted data). Use amultiprocessing.Lockto guard access:def safe_modify(arr, lock): with lock: arr.buf[:] = 42.0 - Windows-specific quirk: On Windows, processes are spawned (not forked), so your
SharedNumpyArrayclass definition must live outside theif __name__ == '__main__'block. This avoids reinitializing shared memory accidentally when the module is reimported in the child. - Lifecycle management: The shared memory segment will be automatically cleaned up once all processes holding a reference to
__bufhave exited. You don't need to manually free it unless you want to release it early.
Example to Demonstrate Shared Behavior
Here's a quick test to confirm the shared memory works as expected:
import multiprocessing as mp import numpy as np from multiprocessing.sharedctypes import RawArray class SharedNumpyArray: def __init__(self, shape, dtype=np.float64): self.shape = shape self.dtype = np.dtype(dtype) self.buf_size = self.dtype.itemsize * np.prod(shape) self.__buf = RawArray('b', self.buf_size) self.init_buf() def init_buf(self): self.buf = np.frombuffer(self.__buf, dtype=self.dtype).reshape(self.shape) def modify_shared_array(arr): # Modify the shared buffer directly arr.buf[:] = 42.0 if __name__ == '__main__': # Create shared array in parent shared_arr = SharedNumpyArray((2, 3)) print("Parent before child modification:\n", shared_arr.buf) # Spawn child process to modify the array p = mp.Process(target=modify_shared_array, args=(shared_arr,)) p.start() p.join() # Check if parent sees the changes print("Parent after child modification:\n", shared_arr.buf)
Running this will output:
Parent before child modification: [[0. 0. 0.] [0. 0. 0.]] Parent after child modification: [[42. 42. 42.] [42. 42. 42.]]
内容的提问来源于stack exchange,提问作者Kiuhnm

