You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Python多进程中执行代码远慢于主进程?

Why is multiprocessing Process much slower than the main process for identical code?

Great question! Let's break down why you're seeing such a massive performance gap between the main process and your spawned child process. The key culprits relate to how multiprocessing handles shared memory, data serialization, and memory access patterns.

1. Shared Memory Array Overhead

The biggest performance drain here is your use of multiprocessing.Array for new_pattern_for_processing:

  • In the main process, you pass a regular Python list to iterate_and_add. Accessing list elements is a fast, direct memory operation with almost no overhead.
  • In the child process, new_pattern_on_new_proc is a shared memory array. Every time you call new_pattern_on_new_proc[count] in your inner loop, you’re not just accessing memory—you’re going through the multiprocessing layer’s synchronization and memory mapping logic.

Your iterate_and_add function runs 2000 * 10 = 20,000 element accesses per pattern. Multiply that by 100 patterns, and all those tiny overheads stack up to the ~17x slowdown you’re seeing in your results.

2. Memory Layout and Cache Efficiency

Another factor is how old_patterns is handled in the child process:

  • In the main process, old_patterns is created sequentially in memory, so the CPU’s cache can efficiently prefetch data as you loop through patterns. This leads to high cache hit rates and fast execution.
  • When you pass old_patterns to the child process via args, it gets serialized (pickled) and deserialized in the child. The deserialized list won’t have the same contiguous memory layout as the original, leading to lower cache hit rates. This makes looping through old_patterns slower in the child.

On Windows (which uses the spawn start method by default), the child even re-executes the module code to initialize itself—though you’re passing old_patterns as an argument, deserialization still disrupts the optimal memory layout.

3. Minor Synchronization Overhead

While less impactful than the first two points, using there_is_new_pattern.Value adds small but cumulative overhead. Every check or update to this value requires acquiring/releasing a lock for process-safe access. Multiply that by 100 patterns, and it contributes to the total slowdown.

How to Fix the Performance Gap

If you need to keep using multiprocessing, these tweaks will bring the child process’s speed closer to the main process:

  • Avoid shared memory for frequent access: Instead of multiprocessing.Array, send entire patterns to the child via a Queue (or Pipe). This gives the child a regular list, eliminating shared memory access overhead.
  • Batch processing: Send multiple patterns at once instead of one by one. This reduces synchronization events and amortizes serialization costs.
  • Use fork (Unix-only): On Linux/macOS, use multiprocessing.set_start_method('fork') at the start of __main__. This clones the main process’s memory directly via copy-on-write, preserving the original memory layout and avoiding serialization overhead.

Example: Using Queue Instead of Shared Memory

Here’s a modified version of your child process code that uses a queue to pass patterns:

def patt_recognition_new_process(old_patterns, queue):
    print_count = 0
    while True:
        new_pattern = queue.get()
        if new_pattern is None:  # Signal to exit
            break
        # Identical code to main process
        start_new_process_one_patt = time.time()
        iterate_and_add(old_patterns, new_pattern)
        if print_count < 10:
            print_count += 1
            print("Time on new process one pattern:", time.time() - start_new_process_one_patt)
        queue.put("DONE")

if __name__ == "__main__":
    # ... main process code ...

    start_new_process = time.time()
    queue = multiprocessing.Queue()
    p1 = multiprocessing.Process(target=patt_recognition_new_process, args=(old_patterns, queue))
    p1.start()
    for new_pattern in new_patterns:
        queue.put(new_pattern)
        while True:
            msg = queue.get()
            if msg == "DONE":
                break
    queue.put(None)  # Tell child process to exit
    p1.join()
    print("Total Time on new process:", time.time()-start_new_process)

This change should drastically reduce the performance gap, as the child now uses regular lists instead of shared memory arrays.

内容的提问来源于stack exchange,提问作者Domen Jakofčič

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 10:00:05