Lisp CFFI调用OpenBLAS gemm较直接C调用慢2倍的原因排查
OpenBLAS dgemm调用性能差异排查:C vs SBCL/CFFI vs MagicL
在Windows 11、MSYS2 MinGW64环境下,使用SBCL 2.5.2调用OpenBLAS 0.3.30执行2000x2000的dgemm运算时,出现显著性能差异:
- 直接C调用仅需0.079秒
- 自行实现的Lisp/CFFI绑定调用耗时0.153秒,性能相差一倍
- 同环境下MagicL库调用相同BLAS的速度与C一致,耗时0.078秒
已确认以下前提:
- 所有调用使用同一OpenBLAS库
- 编译时采用相同优化旗标:
-O3 -mavx -mavx2 -mfma,且添加了-fopenmp - 内存对齐检查无误
- OpenBLAS已启用多线程
快速C版本代码
#include "C:/Users/lisps/nnl2/src/c/nnl2_core.c" #include <time.h> int main() { nnl2_init_system(); Tensor* a = nnl2_ones((int[]){2000, 2000}, 2, FLOAT64); // shape, ndims, dtype Tensor* b = nnl2_ones((int[]){2000, 2000}, 2, FLOAT64); clock_t start = clock(); Tensor* c = gemm(nnl2RowMajor, nnl2Trans, nnl2Trans, 2000, 2000, 2000, 1.0, a, 2000, b, 2000, 0.0); clock_t end = clock(); printf("Time: %f seconds\n", (double)(end - start) / CLOCKS_PER_SEC); // Time: 0.080000 seconds nnl2_free_tensor(a); nnl2_free_tensor(b); nnl2_free_tensor(c); return 0; }
慢速Lisp/CFFI版本代码
(ql:quickload :nnl2) (nnl2.hli.ts:tlet ((a (nnl2.hli.ts:ones #(2000 2000) :dtype :float64)) (b (nnl2.hli.ts:ones #(2000 2000) :dtype :float64))) (time (nnl2.hli.ts:tlet ((c (nnl2.hli.ts:gemm a b)))))) ;; 0.151 seconds (???)
快速MagicL库调用代码
(ql:quickload :magicl) (let ((a (magicl:ones '(2000 2000) :type 'double-float)) (b (magicl:ones '(2000 2000) :type 'double-float))) (time (let ((c (magicl:@ a b)))))) ;; 0.078 seconds
内容的提问来源于stack exchange,提问作者user31676144
相关产品推荐
相关产品推荐

