Mac M1芯片环境下NumPy运算速度异常缓慢问题排查
M1 Mac本地NumPy性能异常问题
问题表现
本地部署的NumPy基础性能测速结果显著低于同场景正常水平:
- 第一轮测试场景为1000×1000随机矩阵乘法,测试代码如下:
import numpy as np A = np.random.rand(1000, 1000) B = np.random.rand(1000, 1000) %timeit A.dot(B)
测试输出结果:
30.3 ms ± 829 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)
同硬件、同场景下矩阵乘法平均耗时普遍低于10ms,本次结果存在明显性能偏差。
环境信息
当前运行环境配置如下:
- 硬件:M1芯片Mac
- 系统:MacOS Big Sur
- Python版本:3.8.13
- NumPy版本:1.22.4,通过
pip install "numpy==1.22.4"命令安装
执行np.show_config()获取的NumPy编译配置信息如下:
openblas64__info: libraries = ['openblas64_', 'openblas64_'] library_dirs = ['/usr/local/lib'] language = c define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None)] runtime_library_dirs = ['/usr/local/lib'] blas_ilp64_opt_info: libraries = ['openblas64_', 'openblas64_'] library_dirs = ['/usr/local/lib'] language = c define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None)] runtime_library_dirs = ['/usr/local/lib'] openblas64__lapack_info: libraries = ['openblas64_', 'openblas64_'] library_dirs = ['/usr/local/lib'] language = c define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None), ('HAVE_LAPACKE', None)] runtime_library_dirs = ['/usr/local/lib'] lapack_ilp64_opt_info: libraries = ['openblas64_', 'openblas64_'] library_dirs = ['/usr/local/lib'] language = c define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None), ('HAVE_LAPACKE', None)] runtime_library_dirs = ['/usr/local/lib'] Supported SIMD extensions in this NumPy install: baseline = SSE,SSE2,SSE3 found = SSSE3,SSE41,POPCNT,SSE42 not found = AVX,F16C,FMA3,AVX2,AVX512F,AVX512CD,AVX512_KNL,AVX512_SKX,AVX512_CLX,AVX512_CNL,AVX512_ICL
补充验证测试
参考公开测试用例开展第二轮性能验证,测试代码如下:
import time import numpy as np np.random.seed(42) a = np.random.uniform(size=(300, 300)) runtimes = 10 timecosts = [] for _ in range(runtimes): s_time = time.time() for i in range(100): a += 1 np.linalg.svd(a) timecosts.append(time.time() - s_time) print(f'mean of {runtimes} runs: {np.mean(timecosts):.5f}s')
本次测试输出结果:
mean of 10 runs: 6.17438s
公开场景下M1 Max芯片不同安装方式的参考测试结果如下:
+-----------------------------------+-----------------------+--------------------+ | Python installed by (run on)→ | Miniforge (native M1) | Anaconda (Rosseta) | +----------------------+------------+------------+----------+----------+---------+ | Numpy installed by ↓ | Run from → | Terminal | PyCharm | Terminal | PyCharm | +----------------------+------------+------------+----------+----------+---------+ | Apple Tensorflow | 4.19151 | 4.86248 | / | / | +-----------------------------------+------------+----------+----------+---------+ | conda install numpy | 4.29386 | 4.98370 | 4.10029 | 4.99271 | +-----------------------------------+------------+----------+----------+---------+
对比可见本次测试耗时显著高于表中所有参考场景,确认存在明确性能异常。
问题根因
从配置信息可直接定位两个核心诱因:
- 安装的NumPy为x86_64架构版本,链接的是
/usr/local/lib路径下的x86版OpenBLAS库,在M1芯片上需要通过Rosetta 2转译运行,转译过程带来大量额外性能开销 - 该NumPy版本编译时仅启用了x86架构的SSE系列SIMD指令集,既没有支持M1芯片专属的NEON、AMX等硬件向量加速指令,连x86架构下的AVX、FMA等高阶计算指令集也未启用,数值计算效率被严重限制
修复方案
- 先卸载当前pip安装的NumPy版本,执行命令:
pip uninstall numpy - 优先安装适配M1芯片ARM64原生架构的Miniforge环境,再通过
conda install numpy命令安装NumPy,该版本会自动链接适配ARM64架构的加速库,可充分调用M1芯片的硬件计算能力 - 若需继续使用pip安装NumPy,需先确认当前运行的Python为ARM64原生版本,避免在x86转译的Python环境中安装x86架构的NumPy包
内容的提问来源于stack exchange,提问作者Jingyang Wang
相关产品推荐
相关产品推荐

