如何在Fortran Coarray中实现异步调用以重叠通信与计算?
默认的Fortran Coarray put操作是阻塞的——它会等数据完全传输完成才返回,没法像MPI_ISEND/MPI_IRECV那样实现通信与计算的重叠。要复刻MPI_I*系列调用的异步能力,你需要用到Fortran 2018标准引入的异步Coarray操作,通过async子句发起后台通信,再用wait语句显式等待操作完成,从而实现通信和计算的并行。
1. 异步Coarray操作的基础用法
发起异步put/get时,只需在Coarray引用后添加async子句,指定一个整数型的请求标识(类似MPI的request句柄):
- 异步Put(把本地数据写入远程图像的Coarray):
! 格式:远程Coarray = 本地数据, async=请求ID neighbours(remote_image)[remote_image]%recv_buff(1:n_send) = neighbours(this_image())%send_buff(1:n_send), async=req_id
- 异步Get(把远程图像的Coarray数据读取到本地):
! 格式:本地数据 = 远程Coarray, async=请求ID neighbours(this_image())%recv_buff(1:n_recv) = neighbours(remote_image)[remote_image]%send_buff(1:n_recv), async=req_id
2. 显式等待异步操作完成
发起异步通信后,你可以立即执行计算任务,之后用wait语句等待指定的请求完成:
wait(req_id) ! 等待单个异步操作完成 wait(req_list) ! req_list是整数数组,等待多个操作完成
3. 适配CFD粒子法的完整示例
针对你的halo粒子交换场景,修改后的异步Coarray代码如下(只和相邻图像通信,更贴合实际需求):
integer, allocatable :: reqs(:) integer :: num_neighbours, i, remote_image, req_idx ! 假设neighbours数组存储当前图像的所有邻居信息,num_neighbours为邻居数量 num_neighbours = size(neighbours) allocate(reqs(2*num_neighbours)) ! 每个邻居对应一个put和一个get请求 req_idx = 0 ! 发起异步Put:把本地待发送的halo粒子写入邻居的接收缓冲区 do i = 1, num_neighbours remote_image = neighbours(i)%image_id if (remote_image /= this_image()) then req_idx = req_idx + 1 neighbours(remote_image)[remote_image]%recv_buff(1:neighbours(i)%n_send) & = neighbours(this_image())%send_buff(1:neighbours(i)%n_send), async=reqs(req_idx) end if end do ! 发起异步Get:从邻居的发送缓冲区读取其halo粒子到本地接收缓冲区 do i = 1, num_neighbours remote_image = neighbours(i)%image_id if (remote_image /= this_image()) then req_idx = req_idx + 1 neighbours(this_image())%recv_buff(1:neighbours(i)%n_recv) & = neighbours(remote_image)[remote_image]%send_buff(1:neighbours(i)%n_recv), async=reqs(req_idx) end if end do ! 执行耗时的计算任务(和后台通信并行) call heavy_computation() ! 等待所有异步通信操作完成 wait(reqs) ! 后续处理halo粒子、推进时间步...
4. 关键注意事项
- 编译器支持:异步Coarray是Fortran 2018的特性,需要使用支持该标准的编译器,比如GCC 10+、Intel Fortran 19+、NVIDIA HPC SDK。
- 数据安全:在异步put操作完成前,不要修改本地的发送缓冲区;在异步get操作完成前,不要读取本地的接收缓冲区,避免数据竞争。
- 请求ID管理:每个异步操作需要唯一的请求标识,确保不会重复使用未完成的ID。
- 旧版本兼容方案:如果编译器不支持Fortran 2018异步特性,可以尝试结合
lock/unlock实现非阻塞通信,但复杂度会接近MPI。
内容的提问来源于stack exchange,提问作者Edward Yang
相关产品推荐
相关产品推荐

