为何numpy.linalg.solve未利用全部可用线程?如何优化性能?
问题:numpy.linalg.solve未充分利用全部可用线程求解多维线性系统
- 需求:重复
n*N*N次求解27维向量的线性系统,预期能充分利用所有线程 - 异常现象:实际仅使用16线程,修改
n(测试300、400)无变化;用threadpoolctl手动指定线程数仅能减少线程数,且无性能提升
测试代码
import numpy as np from numpy import * from timeit import default_timer as timer import math; import os, psutil process = psutil.Process(os.getpid()) ###### Main time loop ########################################################## n = 400 N = 300 s = timer() k_star = np.random.rand(27*n*N*N).reshape((27,n,N,N)) T = np.random.rand(27*27*n*N*N).reshape((27,27,n,N,N)) t1 = timer() - s print("time to allocate k_star and T = " + str(t1)) for time in range(0,1001): s = timer() fout = np.transpose(linalg.solve(np.transpose(T),np.transpose(k_star))) t = timer() - s print("time to calculate fout solving the system = " + str(t)) print(time) if time%1000 == 0: memory = process.memory_info().rss/1000000000 print("RAM Memory occupied = " + str(round(memory,2)) + " Gb")
NumPy配置信息(np.__config__.show()输出)
openblas64__info: libraries = ['openblas64_', 'openblas64_'] library_dirs = ['/usr/local/lib'] language = c define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None)] runtime_library_dirs = ['/usr/local/lib'] blas_ilp64_opt_info: libraries = ['openblas64_', 'openblas64_'] library_dirs = ['/usr/local/lib'] language = c define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None)] runtime_library_dirs = ['/usr/local/lib'] openblas64__lapack_info: libraries = ['openblas64_', 'openblas64_'] library_dirs = ['/usr/local/lib'] language = c define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None), ('HAVE_LAPACKE', None)] runtime_library_dirs = ['/usr/local/lib'] lapack_ilp64_opt_info: libraries = ['openblas64_', 'openblas64_'] library_dirs = ['/usr/local/lib'] language = c define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None), ('HAVE_LAPACKE', None)] runtime_library_dirs = ['/usr/local/lib'] Supported SIMD extensions in this NumPy install: baseline = SSE,SSE2,SSE3 found = SSSE3,SSE41,POPCNT,SSE42,AVX,F16C,FMA3,AVX2,AVX512F,AVX512CD,AVX512_SKX,AVX512_CLX not found = AVX512_KNL,AVX512_KNM,AVX512_CNL,AVX512_ICL
原因分析与解决方法
原因
- OpenBLAS线程限制:从配置看你使用的是OpenBLAS,它的默认线程数可能被环境变量或编译参数固定为16。同时,OpenBLAS对小规模矩阵(27x27)有并行阈值——小矩阵并行的开销大于收益,因此不会自动启用全部线程。
- 批量求解的并行粒度问题:当前写法是一次性传入多维数组让
linalg.solve批量处理,但numpy依赖底层BLAS库的并行逻辑,而BLAS的并行优化更针对单个大矩阵运算,对大量小矩阵的批量任务无法有效扩展到全部线程。
解决方法
调整OpenBLAS线程数
- 先检查环境变量:运行脚本前设置
OPENBLAS_NUM_THREADS为你的CPU核心数,例如:export OPENBLAS_NUM_THREADS=32 # 替换为你的实际核心数 python your_script.py - 如果是编译时固定了线程数,需重新编译OpenBLAS并指定
NUM_THREADS参数,或更换支持动态线程数的预编译版本。
- 先检查环境变量:运行脚本前设置
手动拆分任务,用多进程并行
由于每个27x27的求解任务完全独立,可拆分任务后用进程池并行处理,避开BLAS的小矩阵并行限制:import numpy as np from concurrent.futures import ProcessPoolExecutor def solve_single(T_slice, k_slice): return np.linalg.solve(T_slice, k_slice) # 重构数组,拆分出所有独立的线性系统 T_reshaped = T.transpose(2,3,4,0,1).reshape(-1,27,27) k_reshaped = k_star.transpose(1,2,3,0).reshape(-1,27) # 用进程池并行求解 with ProcessPoolExecutor(max_workers=32) as executor: # 替换为你的核心数 results = list(executor.map(solve_single, T_reshaped, k_reshaped)) # 还原结果到原数组形状 fout = np.array(results).reshape(n,N,N,27).transpose(3,0,1,2)注:使用多进程而非多线程,避免Python GIL对CPU密集型任务的限制。
调整OpenBLAS小矩阵并行阈值
OpenBLAS可通过环境变量或编译参数调整小矩阵的并行策略,例如设置OPENBLAS_CORETYPE或编译时启用USE_SMALL_MATRICS选项,具体可参考OpenBLAS官方文档。
内容的提问来源于stack exchange,提问作者Bruno Magacho
相关产品推荐
相关产品推荐

