MPI_REDUCE部分数组位置数值异常问题求助
MPI_REDUCE后数组出现NaN/超大数值的解决方法
问题场景
使用mpiifort编译器编译以下Fortran MPI代码,调用MPI_REDUCE执行求和操作后,打印recv_results数组时发现:部分位置显示NaN或1e186量级的超大数值,其余位置数值与预期的rad一致。
原代码
program main_mpi_test use mpi implicit none integer(kind=8) :: n integer(kind=8) :: max_optical_depth integer(kind=8) :: bin integer(kind=8) :: tmax integer(kind=8) :: sum_scatt integer(kind=8) :: i,j,k real(kind=8) :: rad real(kind=8), dimension(:), allocatable :: send_results, recv_results integer :: nrank, nproc,ierr,root call MPI_INIT(ierr) call MPI_COMM_RANK(MPI_COMM_WORLD, nrank, ierr) call MPI_COMM_SIZE(MPI_COMM_WORLD, nproc, ierr) call MPI_BARRIER(MPI_COMM_WORLD, ierr) root = 0 max_optical_depth = 10 bin = 10 tmax = max_optical_depth*bin allocate(send_results(tmax)) allocate(recv_results(tmax)) do i = nrank+1, tmax, nproc rad = real(i)/real(bin) send_results(i) = rad print*,'send', rad, send_results(i) end do call MPI_BARRIER(MPI_COMM_WORLD, ierr) call MPI_REDUCE(send_results, recv_results, tmax, MPI_DOUBLE_PRECISION, & MPI_SUM, root, MPI_COMM_WORLD, ierr) if (nrank ==0) then do i = 1, tmax rad = real(i)/real(bin) print*,'recv',rad, recv_results(i) end do end if call MPI_FINALIZE(ierr) deallocate(send_results) deallocate(recv_results) end program main_mpi_test
问题原因
每个进程仅对send_results的部分索引(i = nrank+1, tmax, nproc)赋值,未被赋值的数组元素保留了内存中的随机垃圾值(可能是NaN、极大数或其他无效数据)。而MPI_REDUCE的MPI_SUM操作会对数组的所有tmax个元素执行全局求和,这些垃圾值被累加后,就导致recv_results对应位置出现异常数值。
解决方法
在给send_results赋值前,先将整个数组初始化为0.0d0,确保未被当前进程处理的位置都是合法的0值,这样求和后不会引入垃圾数据。
完整修改后代码
program main_mpi_test use mpi implicit none integer(kind=8) :: n integer(kind=8) :: max_optical_depth integer(kind=8) :: bin integer(kind=8) :: tmax integer(kind=8) :: sum_scatt integer(kind=8) :: i,j,k real(kind=8) :: rad real(kind=8), dimension(:), allocatable :: send_results, recv_results integer :: nrank, nproc,ierr,root call MPI_INIT(ierr) call MPI_COMM_RANK(MPI_COMM_WORLD, nrank, ierr) call MPI_COMM_SIZE(MPI_COMM_WORLD, nproc, ierr) call MPI_BARRIER(MPI_COMM_WORLD, ierr) root = 0 max_optical_depth = 10 bin = 10 tmax = max_optical_depth*bin allocate(send_results(tmax)) allocate(recv_results(tmax)) -- 初始化数组为0,避免垃圾值参与求和 -- send_results = 0.0d0 ------------------------------------ do i = nrank+1, tmax, nproc rad = real(i)/real(bin) send_results(i) = rad print*,'send', rad, send_results(i) end do call MPI_BARRIER(MPI_COMM_WORLD, ierr) call MPI_REDUCE(send_results, recv_results, tmax, MPI_DOUBLE_PRECISION, & MPI_SUM, root, MPI_COMM_WORLD, ierr) if (nrank ==0) then do i = 1, tmax rad = real(i)/real(bin) print*,'recv',rad, recv_results(i) end do end if call MPI_FINALIZE(ierr) deallocate(send_results) deallocate(recv_results) end program main_mpi_test
补充说明
初始化后,每个进程仅在负责的位置写入目标值,其余位置保持0。MPI_REDUCE求和时,每个位置会被唯一的进程贡献目标值,其他进程贡献0,最终recv_results的每个位置都会得到预期的rad值。
内容的提问来源于stack exchange,提问作者HyungJoe Kwon
相关产品推荐
相关产品推荐

