OpenMP矩阵乘法代码:nvfortran/gfortran较Intel慢10倍的原因排查
为何OpenMP Fortran矩阵乘法在Intel编译器下比gfortran/nvfortran快10倍?
我无法理解为何以下OpenMP Fortran矩阵乘法代码,使用nvfortran或gfortran编译后,运行速度比Intel编译器慢约10倍。
测试环境加载命令
module load intel nvhpc/21.9 gcc/11.1.0
测试结果
Intel编译器测试结果
[ilkhom@t019 ORIG]$ bash ./compiler.sh intel The code is compiled with Intel Number of threads | Time (sec) 1 | 2.33 2 | 1.23 3 | 0.85 4 | 0.67 5 | 0.56 6 | 0.46 7 | 0.44 8 | 0.38 9 | 0.37 10 | 0.32 11 | 0.32 12 | 0.28 13 | 0.27 14 | 0.26 15 | 0.24 16 | 0.23
gfortran编译器测试结果
[ilkhom@t019 ORIG]$ bash ./compiler.sh gfortran The code is compiled with gfortran Number of threads | Time (sec) 1 | 31.60 2 | 15.89 3 | 11.33 4 | 8.04 5 | 7.01 6 | 5.65 7 | 4.97 8 | 4.33 9 | 4.15 10 | 3.74 11 | 3.54 12 | 3.11 13 | 2.96 14 | 2.67 15 | 2.59 16 | 2.35
nvfortran编译器测试结果
[ilkhom@t019 ORIG]$ bash ./compiler.sh nvfortran The code is compiled with nvfortran Number of threads | Time (sec) 1 | 31.47 2 | 15.75 3 | 10.73 4 | 8.03 5 | 6.87 6 | 5.67 7 | 4.99 8 | 4.33 9 | 4.19 10 | 3.71 11 | 3.40 12 | 3.10 13 | 3.03 14 | 2.66 15 | 2.75 16 | 2.34
编译脚本 compiler.sh
#! usr/bin/bash if [ $1 == 'intel' ] then ifort -O3 -qopenmp -o matrix_multiply mm.f90 echo "The code is compiled with Intel" elif [ $1 == 'gfortran' ] then gfortran -O3 -fopenmp -o matrix_multiply mm.f90 echo "The code is compiled with gfortran" elif [ $1 == 'nvfortran' ] then nvfortran -O3 -mp -o matrix_multiply mm.f90 echo "The code is compiled with nvfortran" fi echo "Number of threads | Time (sec)" for nthreads in $(seq 1 1 16) do export OMP_NUM_THREADS=$nthreads srun ./matrix_multiply done rm matrix_multiply
Fortran代码 mm.f90
program matrix_multiply use omp_lib implicit none integer, parameter :: N = 3000 integer :: i, j, k real :: start_time, end_time real :: A(N,N), B(N,N), C(N,N) integer :: time_1, time_2, delta_t, countrate, countmax real*8 :: secs ! Initialize matrices A and B do i = 1, N do j = 1, N A(i,j) = i + j B(i,j) = i - j end do end do call system_clock(count_max=countmax, count_rate=countrate) call system_clock(time_1) ! Compute matrix multiplication !$omp parallel do shared(A,B,C) private(i,j,k) do i = 1, N do j = 1, N C(i,j) = 0.0 do k = 1, N C(i,j) = C(i,j) + A(i,k) * B(k,j) end do end do end do !$omp end parallel do call system_clock(time_2) delta_t = time_2-time_1 secs = real(delta_t)/real(countrate) write(6,'(1I5,13X,1A,F10.2)')OMP_GET_MAX_THREADS(),'|',secs end program matrix_multiply
补充说明
gfortran与nvfortran的性能表现相当。我尝试过循环折叠,也显式添加了-Mvect=simd选项(尽管-O3应该已包含该选项),但结果没有变化,性能依旧。Intel编译器显然做了某种优化,我无法理解其中原理,恳请解惑。
内容的提问来源于stack exchange,提问作者Ilkhom Abdurakhmanov
相关产品推荐
相关产品推荐

