You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何numpy.linalg.solve未利用全部可用线程?如何优化性能?

问题:numpy.linalg.solve未充分利用全部可用线程求解多维线性系统
  • 需求:重复n*N*N次求解27维向量的线性系统,预期能充分利用所有线程
  • 异常现象:实际仅使用16线程,修改n(测试300、400)无变化;用threadpoolctl手动指定线程数仅能减少线程数,且无性能提升

测试代码

import numpy as np
from numpy import *
from timeit import default_timer as timer
import math;

import os, psutil
process = psutil.Process(os.getpid())

###### Main time loop ##########################################################

n = 400
N = 300

s = timer()

k_star = np.random.rand(27*n*N*N).reshape((27,n,N,N))
T = np.random.rand(27*27*n*N*N).reshape((27,27,n,N,N))

t1 = timer() - s

print("time to allocate k_star and T = " + str(t1))

for time in range(0,1001):
    s = timer()
    fout = np.transpose(linalg.solve(np.transpose(T),np.transpose(k_star)))
    t = timer() - s
    print("time to calculate fout solving the system = " + str(t))
    print(time)
    
    if time%1000 == 0:
        memory = process.memory_info().rss/1000000000
        print("RAM Memory occupied = " + str(round(memory,2)) + " Gb") 

NumPy配置信息(np.__config__.show()输出)

openblas64__info:
    libraries = ['openblas64_', 'openblas64_']
    library_dirs = ['/usr/local/lib']
    language = c
    define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None)]
    runtime_library_dirs = ['/usr/local/lib']
blas_ilp64_opt_info:
    libraries = ['openblas64_', 'openblas64_']
    library_dirs = ['/usr/local/lib']
    language = c
    define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None)]
    runtime_library_dirs = ['/usr/local/lib']
openblas64__lapack_info:
    libraries = ['openblas64_', 'openblas64_']
    library_dirs = ['/usr/local/lib']
    language = c
    define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None), ('HAVE_LAPACKE', None)]
    runtime_library_dirs = ['/usr/local/lib']
lapack_ilp64_opt_info:
    libraries = ['openblas64_', 'openblas64_']
    library_dirs = ['/usr/local/lib']
    language = c
    define_macros = [('HAVE_CBLAS', None), ('BLAS_SYMBOL_SUFFIX', '64_'), ('HAVE_BLAS_ILP64', None), ('HAVE_LAPACKE', None)]
    runtime_library_dirs = ['/usr/local/lib']
Supported SIMD extensions in this NumPy install:
    baseline = SSE,SSE2,SSE3
    found = SSSE3,SSE41,POPCNT,SSE42,AVX,F16C,FMA3,AVX2,AVX512F,AVX512CD,AVX512_SKX,AVX512_CLX
    not found = AVX512_KNL,AVX512_KNM,AVX512_CNL,AVX512_ICL

原因分析与解决方法

原因

  1. OpenBLAS线程限制:从配置看你使用的是OpenBLAS,它的默认线程数可能被环境变量或编译参数固定为16。同时,OpenBLAS对小规模矩阵(27x27)有并行阈值——小矩阵并行的开销大于收益,因此不会自动启用全部线程。
  2. 批量求解的并行粒度问题:当前写法是一次性传入多维数组让linalg.solve批量处理,但numpy依赖底层BLAS库的并行逻辑,而BLAS的并行优化更针对单个大矩阵运算,对大量小矩阵的批量任务无法有效扩展到全部线程。

解决方法

  1. 调整OpenBLAS线程数

    • 先检查环境变量:运行脚本前设置OPENBLAS_NUM_THREADS为你的CPU核心数,例如:
      export OPENBLAS_NUM_THREADS=32  # 替换为你的实际核心数
      python your_script.py
      
    • 如果是编译时固定了线程数,需重新编译OpenBLAS并指定NUM_THREADS参数,或更换支持动态线程数的预编译版本。
  2. 手动拆分任务,用多进程并行
    由于每个27x27的求解任务完全独立,可拆分任务后用进程池并行处理,避开BLAS的小矩阵并行限制:

    import numpy as np
    from concurrent.futures import ProcessPoolExecutor
    
    def solve_single(T_slice, k_slice):
        return np.linalg.solve(T_slice, k_slice)
    
    # 重构数组,拆分出所有独立的线性系统
    T_reshaped = T.transpose(2,3,4,0,1).reshape(-1,27,27)
    k_reshaped = k_star.transpose(1,2,3,0).reshape(-1,27)
    
    # 用进程池并行求解
    with ProcessPoolExecutor(max_workers=32) as executor:  # 替换为你的核心数
        results = list(executor.map(solve_single, T_reshaped, k_reshaped))
    
    # 还原结果到原数组形状
    fout = np.array(results).reshape(n,N,N,27).transpose(3,0,1,2)
    

    注:使用多进程而非多线程,避免Python GIL对CPU密集型任务的限制。

  3. 调整OpenBLAS小矩阵并行阈值
    OpenBLAS可通过环境变量或编译参数调整小矩阵的并行策略,例如设置OPENBLAS_CORETYPE或编译时启用USE_SMALL_MATRICS选项,具体可参考OpenBLAS官方文档。

内容的提问来源于stack exchange,提问作者Bruno Magacho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 14:19:57