You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pipe和Process并行生成numpy随机数为何结果重复?

Why does using Process + Pipe generate identical random numbers with NumPy?

Great question! The root cause here boils down to how multiprocessing handles memory inheritance versus how threads share memory, combined with NumPy's random number generator (RNG) behavior.

Let's break down the problem

1. Multiprocessing + Pipe issue

When you create a new Process, it forks a copy of the parent process's entire memory space—including NumPy's RNG state. In your code:

  • You start and join each process one after another, and the parent process never modifies its own RNG state between process launches.
  • That means every child process inherits the exact same initial RNG state from the parent. So when each child runs np.random.randint, they're all starting from the same point in the random sequence, hence producing identical results.

2. ThreadPool works as expected

Threads run within the same process, so they share the parent's memory (including the RNG state). Each time you call createNumbers via the thread pool:

  • np.random.randint modifies the shared RNG state after generating numbers.
  • The next call uses the updated state, so it produces a different sequence of random numbers.

Fixes for the Process + Pipe code

You have a few straightforward ways to fix this:

NumPy's modern default_rng() API creates an independent RNG instance, avoiding shared state issues. Use this inside your child function:

import numpy as np
from multiprocessing import Process, Pipe

def createNumbers(conn):
    # Create a fresh, independent RNG for each child process
    rng = np.random.default_rng()
    result = rng.integers(0, 2, size=(4))
    conn.send(result)

parent_conn, child_conn = Pipe()
for i in range(3):
    process = Process(target=createNumbers, args=(child_conn,))
    process.start()
    process.join()
    result = parent_conn.recv()
    print(result)

Option 2: Seed the RNG with a unique value per child

If you prefer using the legacy np.random API, you can seed it with a unique value (like the child process ID) to ensure each process starts with a different state:

import numpy as np
import os
from multiprocessing import Process, Pipe

def createNumbers(conn):
    # Use the child process ID as a unique seed
    np.random.seed(os.getpid())
    result = np.random.randint(0, 2, size=(4))
    conn.send(result)

parent_conn, child_conn = Pipe()
for i in range(3):
    process = Process(target=createNumbers, args=(child_conn,))
    process.start()
    process.join()
    result = parent_conn.recv()
    print(result)

Option 3: Use multiprocessing.Pool (simpler alternative)

If you don't need to use Pipe directly, multiprocessing.Pool handles RNG seeding automatically for you, similar to how ThreadPool works:

import numpy as np
from multiprocessing import Pool

def createNumbers():
    return np.random.randint(0, 2, size=4)

with Pool(processes=3) as pool:
    results = pool.map(createNumbers, range(3))
for result in results:
    print(result)

内容的提问来源于stack exchange,提问作者SebastianH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:15:34