You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Fortran Coarray中实现异步调用以重叠通信与计算?

默认的Fortran Coarray put操作是阻塞的——它会等数据完全传输完成才返回,没法像MPI_ISEND/MPI_IRECV那样实现通信与计算的重叠。要复刻MPI_I*系列调用的异步能力,你需要用到Fortran 2018标准引入的异步Coarray操作,通过async子句发起后台通信,再用wait语句显式等待操作完成,从而实现通信和计算的并行。

1. 异步Coarray操作的基础用法

发起异步put/get时,只需在Coarray引用后添加async子句,指定一个整数型的请求标识(类似MPI的request句柄):

  • 异步Put(把本地数据写入远程图像的Coarray):
! 格式:远程Coarray = 本地数据, async=请求ID
neighbours(remote_image)[remote_image]%recv_buff(1:n_send) = neighbours(this_image())%send_buff(1:n_send), async=req_id
  • 异步Get(把远程图像的Coarray数据读取到本地):
! 格式:本地数据 = 远程Coarray, async=请求ID
neighbours(this_image())%recv_buff(1:n_recv) = neighbours(remote_image)[remote_image]%send_buff(1:n_recv), async=req_id

2. 显式等待异步操作完成

发起异步通信后,你可以立即执行计算任务,之后用wait语句等待指定的请求完成:

wait(req_id)  ! 等待单个异步操作完成
wait(req_list)  ! req_list是整数数组,等待多个操作完成

3. 适配CFD粒子法的完整示例

针对你的halo粒子交换场景,修改后的异步Coarray代码如下(只和相邻图像通信,更贴合实际需求):

integer, allocatable :: reqs(:)
integer :: num_neighbours, i, remote_image, req_idx

! 假设neighbours数组存储当前图像的所有邻居信息,num_neighbours为邻居数量
num_neighbours = size(neighbours)
allocate(reqs(2*num_neighbours))  ! 每个邻居对应一个put和一个get请求
req_idx = 0

! 发起异步Put:把本地待发送的halo粒子写入邻居的接收缓冲区
do i = 1, num_neighbours
  remote_image = neighbours(i)%image_id
  if (remote_image /= this_image()) then
    req_idx = req_idx + 1
    neighbours(remote_image)[remote_image]%recv_buff(1:neighbours(i)%n_send) &
        = neighbours(this_image())%send_buff(1:neighbours(i)%n_send), async=reqs(req_idx)
  end if
end do

! 发起异步Get:从邻居的发送缓冲区读取其halo粒子到本地接收缓冲区
do i = 1, num_neighbours
  remote_image = neighbours(i)%image_id
  if (remote_image /= this_image()) then
    req_idx = req_idx + 1
    neighbours(this_image())%recv_buff(1:neighbours(i)%n_recv) &
        = neighbours(remote_image)[remote_image]%send_buff(1:neighbours(i)%n_recv), async=reqs(req_idx)
  end if
end do

! 执行耗时的计算任务(和后台通信并行)
call heavy_computation()

! 等待所有异步通信操作完成
wait(reqs)

! 后续处理halo粒子、推进时间步...

4. 关键注意事项

  • 编译器支持:异步Coarray是Fortran 2018的特性,需要使用支持该标准的编译器,比如GCC 10+、Intel Fortran 19+、NVIDIA HPC SDK。
  • 数据安全:在异步put操作完成前,不要修改本地的发送缓冲区;在异步get操作完成前,不要读取本地的接收缓冲区,避免数据竞争。
  • 请求ID管理:每个异步操作需要唯一的请求标识,确保不会重复使用未完成的ID。
  • 旧版本兼容方案:如果编译器不支持Fortran 2018异步特性,可以尝试结合lock/unlock实现非阻塞通信,但复杂度会接近MPI。

内容的提问来源于stack exchange,提问作者Edward Yang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 11:53:13