You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何预转置矩阵的矩阵乘法运算速度更快?

预转置矩阵相乘为何比非转置更快?

先看初始测试代码:

import numpy as np
import time

# 生成随机矩阵
matrix_size = 1000
matrix = np.random.rand(matrix_size, matrix_size)

# 转置矩阵
transposed_matrix = np.transpose(matrix)

# 非转置矩阵相乘
start = time.time()
result1 = np.matmul(matrix, matrix)
end = time.time()
execution_time1 = end - start

# 预转置矩阵相乘
start = time.time()
result2 = np.matmul(transposed_matrix, transposed_matrix)
end = time.time()
execution_time2 = end - start

print("非转置执行时间:", execution_time1)
print("预转置执行时间:", execution_time2)

运行后发现预转置矩阵相乘的速度明显更快——直觉上乘法顺序不该影响性能,但实际差异显著。为什么会出现这种情况?是否有底层机制可以解释?


更新

考虑到缓存的影响,我修改了代码,每次循环生成新矩阵以排除缓存预热干扰:

import numpy as np
import time
import matplotlib.pyplot as plt

# 矩阵尺寸
matrix_size = 3000

# 存储执行时间的变量
execution_times1 = []
execution_times2 = []

# 执行50次A@A并记录时间
num_iterations = 50
for _ in range(num_iterations):
    matrix_a = np.random.rand(matrix_size, matrix_size)
    start = time.time()
    result1 = np.matmul(matrix_a, matrix_a)
    end = time.time()
    execution_times1.append(end - start)

# 执行50次B@B.T并记录时间
for _ in range(num_iterations):
    matrix_b = np.random.rand(matrix_size, matrix_size)
    start = time.time()
    result2 = np.matmul(matrix_b, matrix_b.T)
    end = time.time()
    execution_times2.append(end - start)

# 计算平均执行时间
avg_execution_time1 = np.mean(execution_times1)
avg_execution_time2 = np.mean(execution_times2)

# 绘制执行时间对比图
plt.plot(range(num_iterations), execution_times1, label='A @ A')
plt.plot(range(num_iterations), execution_times2, label='B @ B.T')
plt.xlabel('迭代次数')
plt.ylabel('执行时间')
plt.title('矩阵乘法执行时间对比')
plt.legend()
plt.show()

# 显示BLAS配置
np.show_config()

结果

执行时间对比图

矩阵乘法执行时间对比

BLAS配置信息

blas_mkl_info:
    libraries = ['mkl_rt']
    library_dirs = ['C:/Users/User/anaconda3\\Library\\lib']
    define_macros = [('SCIPY_MKL_H', None), ('HAVE_CBLAS', None)]
    include_dirs = ['C:/Users/User/anaconda3\\Library\\include']
blas_opt_info:
    libraries = ['mkl_rt']
    library_dirs = ['C:/Users/User/anaconda3\\Library\\lib']
    define_macros = [('SCIPY_MKL_H', None), ('HAVE_CBLAS', None)]
    include_dirs = ['C:/Users/User/anaconda3\\Library\\include']
lapack_mkl_info:
    libraries = ['mkl_rt']
    library_dirs = ['C:/Users/User/anaconda3\\Library\\lib']
    define_macros = [('SCIPY_MKL_H', None), ('HAVE_CBLAS', None)]
    include_dirs = ['C:/Users/User/anaconda3\\Library\\include']
lapack_opt_info:
    libraries = ['mkl_rt']
    library_dirs = ['C:/Users/User/anaconda3\\Library\\lib']
    define_macros = [('SCIPY_MKL_H', None), ('HAVE_CBLAS', None)]
    include_dirs = ['C:/Users/User/anaconda3\\Library\\include']
Supported SIMD extensions in this NumPy install:
    baseline = SSE,SSE2,SSE3
    found = SSSE3,SSE41,POPCNT,SSE42,AVX,F16C,FMA3,AVX2
    not found = AVX512F,AVX512CD,AVX512_SKX,AVX512_CLX,AVX512_CNL

原因解析

这个性能差异的核心在于CPU缓存命中率和BLAS库的底层优化:

  1. 内存访问模式的缓存友好性
    NumPy默认采用**行优先(C风格)**存储矩阵——同一行的元素在内存中是连续存放的。
  • 计算A @ A时,第一个矩阵A按行读取(连续内存,缓存命中率高),但第二个矩阵A需要按列读取(非连续内存,缓存频繁失效)。矩阵乘法中第二个矩阵的列元素会被反复访问,而这些元素在内存中分散存储,需要频繁从主存读取,拖慢计算速度。
  • 计算B @ B.T时,B.T的行对应原矩阵B的列,转置后这些列元素变成连续内存。此时两个矩阵都按行读取,都是连续内存访问,缓存命中率大幅提升,计算速度自然更快。
  1. BLAS库的针对性优化
    从配置可知你使用的是Intel MKL作为BLAS后端,这类高性能库会针对内存布局做深度优化:
  • 它会自动检测矩阵的存储布局,选择最优计算路径。对于转置后的矩阵相乘,MKL能充分利用连续内存的优势,配合AVX2、FMA3等SIMD指令批量处理数据,进一步提升效率。
  • 此外,MKL会采用分块(Blocking)策略,将大矩阵拆分成适配CPU缓存大小的小块,减少缓存失效次数,这在连续内存访问场景下效果更显著。
  1. 后续验证建议
    如果想进一步验证,可以尝试:
  • 使用np.asfortranarray()将矩阵转为列优先存储,对比A @ A和A.T @ A.T的性能差异。
  • 测试不同尺寸的矩阵,观察小矩阵(完全能放进缓存)和大矩阵的性能差异是否一致。

内容的提问来源于stack exchange,提问作者zoldxk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 21:55:20