You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Numpy:省略优化操作与原地操作的性能差异原因咨询

Numpy原地操作与Elision优化的性能差异问题

我了解Numpy支持的elision优化,原本认为以下两种操作均无需分配新内存,性能应一致,但测试发现elided_then_copyto耗时仅为copyto_then_incr的60%。

测试代码

import numpy as np
import cProfile

# 生成图像蓝色通道的随机输入
im_shape = (900, 1600, 3)
foo_B_chan = np.random.randint(low=0, high=256, size=im_shape[:2], dtype=np.uint8)
# 预分配输出缓冲区
y_buf = np.ndarray(im_shape, dtype=np.uint8)
# 单独提取蓝色通道
y_buf_B_chan = y_buf[:,:,2]

def elided_then_copyto(x_B_chan, y_B_chan):
    z = x_B_chan + 1
    np.copyto(y_B_chan, z)

def copyto_then_incr(x_B_chan, y_B_chan):
    np.copyto(y_B_chan, x_B_chan)
    y_B_chan += 1

with cProfile.Profile() as pr:
    for j in np.arange(1000):
        elided_then_copyto(foo_B_chan, y_buf_B_chan)
        copyto_then_incr(foo_B_chan, y_buf_B_chan)
    pr.print_stats(sort=2)

两种方法的字节码符合预期,但通过pprof分析发现,原地操作的加法速度远慢于省略优化操作(前者add_AVX2耗时1996ms,后者仅233ms),甚至数组复制也更慢。请问原地操作的额外开销来源于何处?


内容的提问来源于stack exchange,提问作者stephematician

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 02:48:22