使用OMP if和reduction时gfortran/OpenMP是否存在Bug?
我编写了一个OpenMP最小可复现示例,当保留GPU offloading相关的target teams if(is_GPU)结构(即使is_GPU为.false.,即实际运行CPU模式)时,gfortran编译后的max_diff归约结果异常(输出0.000,预期0.900),但nvfortran编译运行结果完全正常。目前可以通过预处理器指令作为临时解决方案,但希望采用OpenMP原生方案。
测试环境
编译器版本
- gfortran 10.3.0、11.3.0
- nvfortran 22.5-0(来自NVHPC工具包)
硬件
- Arm Neoverse-N1 CPU
- Intel Core i7-7500U CPU
编译命令
- GPU版本(nvfortran):
nvfortran -cpp -DUSEGPU -mp=gpu mwe.f90 && ./a.out - CPU版本(nvfortran):
nvfortran -cpp -mp=multicore mwe.f90 && OMP_NUM_THREADS=2 ./a.out - CPU版本(gfortran):
gfortran -cpp -fopenmp mwe.f90 && OMP_NUM_THREADS=2 ./a.out
问题现象
预期最终max_diff值为0.900,但gfortran编译的CPU模式输出max_diff = 0.000。若注释掉标记为!<!-- comment this -->的!$omp target teams if(is_GPU)及对应结束语句,并删除distribute关键字,gfortran编译运行结果恢复正常。
问题原因
这是gfortran的OpenMP实现Bug:
根据OpenMP标准,当target teams if(condition)中的条件为假时,该构造应降级为主机(CPU)上的普通并行区域,相当于忽略target teams和distribute,直接执行内部的parallel do simd。但gfortran在这种场景下,没有正确处理归约变量max_diff的线程间同步,导致最终结果保留初始值0.0。而nvfortran的实现正确处理了这种降级逻辑,因此CPU模式下结果正常。
解决方案
方案1:OpenMP原生条件分支(推荐)
使用OpenMP的!$omp if指令替代预处理器,在CPU模式下跳过target结构,让gfortran正确处理归约:
max_diff = 0.0 !$omp if (is_GPU) then !$omp target teams !$omp distribute parallel do simd reduction(max:max_diff) collapse(2) !$omp else !$omp parallel do simd reduction(max:max_diff) collapse(2) !$omp end if do j = 1, N do i = 1, N max_diff = max(max_diff, abs(b(i, j) - a(i, j))) #ifndef USEGPU write (*,'("id=",I2," i,j= ",2(I2)," a_ij= ",F5.3," b_ij= ",F5.3," max_diff=",F5.3)') & omp_get_thread_num(), i, j, a(i,j), b(i,j), max_diff #endif end do end do !$omp if (is_GPU) then !$omp end target teams !$omp end if
方案2:升级gfortran版本
gfortran 12.x及以上版本已修复部分OpenMP target降级逻辑的Bug,可尝试升级编译器解决问题。
方案3:预处理器临时方案(已提及)
用#ifdef USEGPU包裹target块,CPU模式下直接使用普通并行循环:
max_diff = 0.0 #ifdef USEGPU !$omp target teams !$omp distribute parallel do simd reduction(max:max_diff) collapse(2) #else !$omp parallel do simd reduction(max:max_diff) collapse(2) #endif do j = 1, N do i = 1, N max_diff = max(max_diff, abs(b(i, j) - a(i, j))) #ifndef USEGPU write (*,'("id=",I2," i,j= ",2(I2)," a_ij= ",F5.3," b_ij= ",F5.3," max_diff=",F5.3)') & omp_get_thread_num(), i, j, a(i,j), b(i,j), max_diff #endif end do end do #ifdef USEGPU !$omp end target teams #endif
原测试代码(mwe.f90)
program test use omp_lib implicit none integer, parameter :: N=3 integer :: i, j real :: a(N,N), b(N,N), max_diff logical :: is_GPU is_GPU = .false. #ifdef USEGPU is_GPU = .true. #endif !$omp target data if(is_GPU) map(to:a, b) !$omp target teams if(is_GPU) !$omp distribute parallel do simd collapse(2) do j = 1, N do i = 1, N a(i, j) = i*j b(i, j) = i*j*0.9 end do end do !$omp end target teams max_diff = 0.0 !$omp target teams if(is_GPU) !<!-- comment this --> !$omp distribute parallel do simd reduction(max:max_diff) collapse(2) do j = 1, N do i = 1, N max_diff = max(max_diff, abs(b(i, j) - a(i, j))) #ifndef USEGPU write (*,'("id=",I2," i,j= ",2(I2)," a_ij= ",F5.3," b_ij= ",F5.3," max_diff=",F5.3)') & omp_get_thread_num(), i, j, a(i,j), b(i,j), max_diff #endif end do end do !$omp end target teams !<!-- comment this --> write (*,'("max_diff = ", F6.3)') max_diff !$omp end target data end program
内容的提问来源于stack exchange,提问作者melt

