You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何10核处理器上Numba可将随机数生成加速约175倍?

问题分析与解答

问题背景

我在两台机器上得到了大致相似的结果:一台搭载Intel i5-1235U(10核,2个性能核,共12线程),另一台搭载Intel i7-13700KF(16核,8个性能核,共24线程)。
我正在学习并行计算,编写了一个简单脚本对比Numba串行与并行任务的耗时,任务为生成n个0-9之间的随机整数。我原本预期最大加速比应为硬件总线程数(12倍或24倍),但实际观测到最高达175倍的性能提升。若线程数并非限制因素,那原因是什么?额外性能是否来自超线程?

测试代码

import numpy as np
from numba import jit, prange
import time

@jit(nopython=True, parallel=True)
def generate_random_vector_parallel(n):
    vector = np.empty(n, dtype=np.uint8)
    for i in prange(n):
        vector[i] = np.random.randint(0, 10)
    return vector


def generate_random_vector_serial(n):
    vector = np.empty(n, dtype=np.uint8)
    for i in range(n):
        vector[i] = np.random.randint(0, 10)
    return vector

n = 10000
n_linspace = range(1, n + 1)
parallel_times = np.zeros(n)
serial_times = np.zeros(n)


def iterate(i):
    parallel_start_time = time.time()
    parallel_random_vector = generate_random_vector_parallel(i)
    parallel_time_difference = time.time() - parallel_start_time
    parallel_times[i - 1] += parallel_time_difference

    serial_start_time = time.time()
    serial_random_vector = generate_random_vector_serial(i)
    serial_time_difference = time.time() - serial_start_time
    serial_times[i - 1] += serial_time_difference

superiterations = 10

for m in range(superiterations):
    for n in n_linspace:
        iterate(n)

parallel_times = parallel_times / superiterations
serial_times = serial_times / superiterations

serial_over_parallel = [
    serial / parallel for serial, parallel in zip(serial_times, parallel_times)
]



# Block to plot with matplotlib
import matplotlib.backends.backend_pdf

pdf = matplotlib.backends.backend_pdf.PdfPages("output.pdf")

fig1 = plt.figure()
plt.title("Average execution time of random vector of length\n(parallel) (n=100)")
plt.ylabel("Average execution time (s)")
plt.xlabel("Length of vector")
plt.plot(n_linspace[1:], parallel_times[1:], color="blue")
plt.show()

fig2 = plt.figure()
plt.title("Average execution time of random vector of length\n(serial) (n=100)")
plt.ylabel("Average execution time (s)")
plt.xlabel("Length of vector")
plt.plot(n_linspace[1:], serial_times[1:], color="red")
plt.show()

fig3 = plt.figure()
plt.title("Average execution time of random vector of length\n(parallel) (n=100)")
plt.ylabel("log Average execution time (s)")
plt.xlabel("log Length of vector")
plt.loglog(n_linspace[1:], parallel_times[1:], color="blue")
plt.loglog(n_linspace[1:], serial_times[1:], color="red")
plt.show()

fig4 = plt.figure()
plt.title("Average execution time of random vector of length\n(parallel) (n=100)")
plt.ylabel("(avg serial execution time) /(avg parallel execution time)")
plt.xlabel("Length of vector")
plt.plot(n_linspace[1:], serial_over_parallel[1:], color="grey")
plt.show()

pdf.savefig(fig1)
pdf.savefig(fig2)
pdf.savefig(fig3)
pdf.savefig(fig4)

pdf.close()

测试图表

  • 并行执行时间曲线:
    Parallel speed
  • 串行执行时间曲线:
    Serial speed
  • 双对数坐标下的执行时间对比:
    loglog of both speeds
  • 串行/并行加速比曲线:
    serial speed / parallel speed

核心原因分析

你观测到的远超硬件线程数的加速比,根本原因不是超线程,而是两个版本的代码在底层实现上的差异远不止“串行/并行”的区别:

  1. Numba JIT编译优化 vs 纯Python解释执行
    并行函数用了@jit(nopython=True, parallel=True)装饰器,Numba会把这部分代码编译成机器码直接执行;而串行函数是纯Python循环,完全依赖Python解释器逐行执行——Python解释器的单循环开销本身就极大,这才是两者性能差距的主要来源,而非并行本身。

  2. 随机数生成的实现差异

    • Numba在nopython模式下调用的np.random.randint是经过编译的高效实现,而纯Python循环里调用的是Python层面的np.random.randint,每次调用都要经过解释器的函数调用开销,累加起来差异巨大。
    • 并行版本中,Numba还可能为多线程分配了独立的随机数生成器状态,避免了锁竞争,进一步提升了效率。
  3. 小任务下的测量偏差
    从图表看,小n值时加速比极高——这是因为小任务的绝对执行时间极短,Python解释器的启动、函数调用等固定开销占比被放大,而Numba编译后的代码几乎没有这些额外开销,导致计算出的加速比虚高。当n足够大时,加速比会逐渐收敛到接近硬件线程数的水平(从最后一张图也能看到,随着n增大,加速比趋于稳定)。

  4. 超线程的作用有限
    超线程确实能提升CPU利用率,但最多只能带来约1.5-2倍的单核心性能提升,不可能产生175倍的差距。你的测试结果里,超线程的贡献可以忽略不计,主要性能差距来自JIT编译和Python解释器的本质差异。

验证建议

如果要真正测试并行带来的加速比,应该让两个版本都经过Numba JIT编译,只保留“串行/并行”这一个变量差异:

# 改为两个都用Numba编译,仅循环方式不同
@jit(nopython=True)
def generate_random_vector_serial_jit(n):
    vector = np.empty(n, dtype=np.uint8)
    for i in range(n):
        vector[i] = np.random.randint(0, 10)
    return vector

此时对比generate_random_vector_serial_jit和generate_random_vector_parallel的耗时,得到的加速比才会接近硬件线程数的理论上限。

内容的提问来源于stack exchange,提问作者fngarrett

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 21:42:02