You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenMP矩阵乘法代码:nvfortran/gfortran较Intel慢10倍的原因排查

为何OpenMP Fortran矩阵乘法在Intel编译器下比gfortran/nvfortran快10倍?

我无法理解为何以下OpenMP Fortran矩阵乘法代码,使用nvfortran或gfortran编译后,运行速度比Intel编译器慢约10倍。

测试环境加载命令

module load intel nvhpc/21.9 gcc/11.1.0

测试结果

Intel编译器测试结果

[ilkhom@t019 ORIG]$ bash ./compiler.sh intel
The code is compiled with Intel
Number of threads |   Time (sec)
    1             |      2.33
    2             |      1.23
    3             |      0.85
    4             |      0.67
    5             |      0.56
    6             |      0.46
    7             |      0.44
    8             |      0.38
    9             |      0.37
   10             |      0.32
   11             |      0.32
   12             |      0.28
   13             |      0.27
   14             |      0.26
   15             |      0.24
   16             |      0.23

gfortran编译器测试结果

[ilkhom@t019 ORIG]$ bash ./compiler.sh gfortran
The code is compiled with gfortran
Number of threads |   Time (sec)
    1             |     31.60
    2             |     15.89
    3             |     11.33
    4             |      8.04
    5             |      7.01
    6             |      5.65
    7             |      4.97
    8             |      4.33
    9             |      4.15
   10             |      3.74
   11             |      3.54
   12             |      3.11
   13             |      2.96
   14             |      2.67
   15             |      2.59
   16             |      2.35

nvfortran编译器测试结果

[ilkhom@t019 ORIG]$ bash ./compiler.sh nvfortran
The code is compiled with nvfortran
Number of threads |   Time (sec)
    1             |     31.47
    2             |     15.75
    3             |     10.73
    4             |      8.03
    5             |      6.87
    6             |      5.67
    7             |      4.99
    8             |      4.33
    9             |      4.19
   10             |      3.71
   11             |      3.40
   12             |      3.10
   13             |      3.03
   14             |      2.66
   15             |      2.75
   16             |      2.34

编译脚本 compiler.sh

#! usr/bin/bash

if [ $1 == 'intel' ]
then
  ifort -O3 -qopenmp -o matrix_multiply mm.f90
  echo "The code is compiled with Intel"  
elif [ $1 == 'gfortran' ]
then
  gfortran -O3 -fopenmp -o matrix_multiply mm.f90
  echo "The code is compiled with gfortran"  
elif [ $1 == 'nvfortran' ]
then
  nvfortran -O3 -mp -o matrix_multiply mm.f90
  echo "The code is compiled with nvfortran"  
fi

echo "Number of threads |   Time (sec)"

for nthreads in $(seq 1 1 16)
    do
      export OMP_NUM_THREADS=$nthreads 
      srun ./matrix_multiply
done

rm matrix_multiply

Fortran代码 mm.f90

program matrix_multiply
    use omp_lib
    implicit none
    integer, parameter :: N = 3000
    integer :: i, j, k
    real :: start_time, end_time
    real :: A(N,N), B(N,N), C(N,N)
    integer :: time_1, time_2, delta_t, countrate, countmax
    real*8 :: secs
    
    ! Initialize matrices A and B
    do i = 1, N
        do j = 1, N
            A(i,j) = i + j
            B(i,j) = i - j
        end do
    end do
    
    call system_clock(count_max=countmax, count_rate=countrate)
    call system_clock(time_1)
    ! Compute matrix multiplication
    !$omp parallel do shared(A,B,C) private(i,j,k)
    do i = 1, N
        do j = 1, N
            C(i,j) = 0.0
            do k = 1, N
                C(i,j) = C(i,j) + A(i,k) * B(k,j)
            end do
        end do
    end do
    !$omp end parallel do
    call system_clock(time_2)
    delta_t = time_2-time_1
    secs = real(delta_t)/real(countrate)

    write(6,'(1I5,13X,1A,F10.2)')OMP_GET_MAX_THREADS(),'|',secs
    
end program matrix_multiply

补充说明

gfortran与nvfortran的性能表现相当。我尝试过循环折叠,也显式添加了-Mvect=simd选项(尽管-O3应该已包含该选项),但结果没有变化,性能依旧。Intel编译器显然做了某种优化,我无法理解其中原理,恳请解惑。


内容的提问来源于stack exchange,提问作者Ilkhom Abdurakhmanov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 20:05:07