Numpy技术疑问:内存布局不一致且尺寸为2的大幂次时逐元素运算为何大幅变慢?
Numpy异布局矩阵逐元素运算的性能异常问题
我在测试Numpy运算性能时发现:对内存布局不一致(一个为C序、一个为F序)的矩阵执行逐元素运算(例如np.multiply),比对内存布局匹配的矩阵执行同类运算慢约2倍,这符合内存局部性原理的预期。
但令我惊讶的是,当矩阵尺寸为2的大幂次时,内存布局不一致带来的性能惩罚会显著增大,运算速度比布局匹配的情况慢4倍甚至更多。比如1024x1024的异布局矩阵逐元素乘法耗时,与1800x1800的同布局矩阵运算耗时相近。
硬件配置
以下是lshw -short的输出:
H/W path Device Class Description ============================================================== system 20KN001QPB (LENOVO_MT_20KN_BU_Think_FM_ThinkPad E480) /0 bus 20KN001QPB /0/3 memory 16GiB System Memory /0/3/0 memory 8GiB SODIMM DDR4 Synchronous Unbuffered (Unregistered) 2400 MHz (0,4 ns) /0/3/1 memory 8GiB SODIMM DDR4 Synchronous Unbuffered (Unregistered) 2400 MHz (0,4 ns) /0/7 memory 256KiB L1 cache /0/8 memory 1MiB L2 cache /0/9 memory 6MiB L3 cache /0/a processor Intel(R) Core(TM) i5-8250U CPU @ 1.60GHz /0/b memory 128KiB BIOS /0/100 bridge Xeon E3-1200 v6/7th Gen Core Processor Host Bridge/DRAM Registers /0/100/2 display UHD Graphics 620 /0/100/8 generic Xeon E3-1200 v5/v6 / E3-1500 v5 / 6th/7th/8th ...
实验环境
- Python 3.9.2
- numpy 1.23.4
复现实验代码
import numpy as np from matplotlib import pyplot as plt from timeit import default_timer as timer def measure_time(data, op, rep=10): start = timer() for i in range(rep): op(data) return (timer() - start) / rep def prep_matrix(n, order='C'): return np.ones((n, n), order=order) def mul(data): a, b = data np.multiply(a, b) times = [measure_time((prep_matrix(i), prep_matrix(i)), mul, 3) for i in range(0, 3100, 1)] times2 = [measure_time((prep_matrix(i), prep_matrix(i, order='F')), mul, 3) for i in range(0, 3100, 1)] plt.plot(times) plt.plot(times2) plt.show()
实验结果
实验结果图中,橙色曲线代表内存布局不同的矩阵乘法执行时间(单位:秒),蓝色曲线代表内存布局匹配的情况,X轴对应矩阵尺寸。
显然存在异常现象,请问这是什么原因导致的?
内容的提问来源于stack exchange,提问作者krzysztor
相关产品推荐
相关产品推荐

