为何Python多进程中执行代码远慢于主进程?
Great question! Let's break down why you're seeing such a massive performance gap between the main process and your spawned child process. The key culprits relate to how multiprocessing handles shared memory, data serialization, and memory access patterns.
1. Shared Memory Array Overhead
The biggest performance drain here is your use of multiprocessing.Array for new_pattern_for_processing:
- In the main process, you pass a regular Python list to
iterate_and_add. Accessing list elements is a fast, direct memory operation with almost no overhead. - In the child process,
new_pattern_on_new_procis a shared memory array. Every time you callnew_pattern_on_new_proc[count]in your inner loop, you’re not just accessing memory—you’re going through the multiprocessing layer’s synchronization and memory mapping logic.
Your iterate_and_add function runs 2000 * 10 = 20,000 element accesses per pattern. Multiply that by 100 patterns, and all those tiny overheads stack up to the ~17x slowdown you’re seeing in your results.
2. Memory Layout and Cache Efficiency
Another factor is how old_patterns is handled in the child process:
- In the main process,
old_patternsis created sequentially in memory, so the CPU’s cache can efficiently prefetch data as you loop through patterns. This leads to high cache hit rates and fast execution. - When you pass
old_patternsto the child process viaargs, it gets serialized (pickled) and deserialized in the child. The deserialized list won’t have the same contiguous memory layout as the original, leading to lower cache hit rates. This makes looping throughold_patternsslower in the child.
On Windows (which uses the spawn start method by default), the child even re-executes the module code to initialize itself—though you’re passing old_patterns as an argument, deserialization still disrupts the optimal memory layout.
3. Minor Synchronization Overhead
While less impactful than the first two points, using there_is_new_pattern.Value adds small but cumulative overhead. Every check or update to this value requires acquiring/releasing a lock for process-safe access. Multiply that by 100 patterns, and it contributes to the total slowdown.
How to Fix the Performance Gap
If you need to keep using multiprocessing, these tweaks will bring the child process’s speed closer to the main process:
- Avoid shared memory for frequent access: Instead of
multiprocessing.Array, send entire patterns to the child via aQueue(orPipe). This gives the child a regular list, eliminating shared memory access overhead. - Batch processing: Send multiple patterns at once instead of one by one. This reduces synchronization events and amortizes serialization costs.
- Use
fork(Unix-only): On Linux/macOS, usemultiprocessing.set_start_method('fork')at the start of__main__. This clones the main process’s memory directly via copy-on-write, preserving the original memory layout and avoiding serialization overhead.
Example: Using Queue Instead of Shared Memory
Here’s a modified version of your child process code that uses a queue to pass patterns:
def patt_recognition_new_process(old_patterns, queue): print_count = 0 while True: new_pattern = queue.get() if new_pattern is None: # Signal to exit break # Identical code to main process start_new_process_one_patt = time.time() iterate_and_add(old_patterns, new_pattern) if print_count < 10: print_count += 1 print("Time on new process one pattern:", time.time() - start_new_process_one_patt) queue.put("DONE") if __name__ == "__main__": # ... main process code ... start_new_process = time.time() queue = multiprocessing.Queue() p1 = multiprocessing.Process(target=patt_recognition_new_process, args=(old_patterns, queue)) p1.start() for new_pattern in new_patterns: queue.put(new_pattern) while True: msg = queue.get() if msg == "DONE": break queue.put(None) # Tell child process to exit p1.join() print("Total Time on new process:", time.time()-start_new_process)
This change should drastically reduce the performance gap, as the child now uses regular lists instead of shared memory arrays.
内容的提问来源于stack exchange,提问作者Domen Jakofčič

