You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MPI乒乓代码运行出现空闲停滞问题求助

问题描述

在ArchLinux平台使用OpenMPI 4.1.5运行MPI乒乓代码时,程序在ping_pong_count到4后无后续输出,无法达到设定的PING_PONG_LIMIT=10。

运行环境

  • 操作系统:ArchLinux
  • OpenMPI版本:4.1.5
  • 测试代码:MPI乒乓示例代码

代码内容

// Author: Wes Kendall
// Copyright 2011 www.mpitutorial.com
// This code is provided freely with the tutorials on mpitutorial.com. Feel
// free to modify it for your own use. Any distribution of the code must
// either provide a link to www.mpitutorial.com or keep this header intact.
//
// Ping pong example with MPI_Send and MPI_Recv. Two processes ping pong a
// number back and forth, incrementing it until it reaches a given value.
//
#include <mpi.h>
#include <stdio.h>
#include <stdlib.h>

int main(int argc, char** argv) {
  const int PING_PONG_LIMIT = 10;

  // Initialize the MPI environment
  MPI_Init(NULL, NULL);
  // Find out rank, size
  int world_rank;
  MPI_Comm_rank(MPI_COMM_WORLD, &world_rank);
  int world_size;
  MPI_Comm_size(MPI_COMM_WORLD, &world_size);

  // We are assuming 2 processes for this task
  if (world_size != 2) {
    fprintf(stderr, "World size must be two for %s\n", argv[0]);
    MPI_Abort(MPI_COMM_WORLD, 1);
  }

  int ping_pong_count = 0;
  int partner_rank = (world_rank + 1) % 2;
  while (ping_pong_count < PING_PONG_LIMIT) {
    if (world_rank == ping_pong_count % 2) {
      // Increment the ping pong count before you send it
      ping_pong_count++;
      MPI_Send(&ping_pong_count, 1, MPI_INT, partner_rank, 0, MPI_COMM_WORLD);
      printf("%d sent and incremented ping_pong_count %d to %d\n",
             world_rank, ping_pong_count, partner_rank);
    } else {
      MPI_Recv(&ping_pong_count, 1, MPI_INT, partner_rank, 0, MPI_COMM_WORLD,
               MPI_STATUS_IGNORE);
      printf("%d received ping_pong_count %d from %d\n",
             world_rank, ping_pong_count, partner_rank);
    }
  }
  MPI_Finalize();
}

执行日志

$ mpicc ./ping_pong.c -o ping_pong
$ mpirun -np 2 ./ping_pong        
0 sent and incremented ping_pong_count 1 to 1
0 received ping_pong_count 2 from 1
0 sent and incremented ping_pong_count 3 to 1
0 received ping_pong_count 4 from 1
1 received ping_pong_count 1 from 0
1 sent and incremented ping_pong_count 2 to 0
1 received ping_pong_count 3 from 0
1 sent and incremented ping_pong_count 4 to 0
^C

问题原因与解决方法

核心原因

程序并未真正卡住,而是标准输出的行缓冲机制导致后续输出延迟显示:

  • MPI进程的输出是独立缓冲的,默认printf仅在缓冲区满或遇到换行符时刷新到终端,但在部分MPI运行环境中,进程输出会被延迟合并,导致用户误以为程序停止运行。
  • 从日志顺序可见,进程0的输出先批量显示,之后才是进程1的输出,印证了输出缓冲未及时刷新的问题。

解决方法

在每次printf调用后强制刷新标准输出缓冲区,确保内容立即显示:

修改后的关键代码片段:

if (world_rank == ping_pong_count % 2) {
  ping_pong_count++;
  MPI_Send(&ping_pong_count, 1, MPI_INT, partner_rank, 0, MPI_COMM_WORLD);
  printf("%d sent and incremented ping_pong_count %d to %d\n",
         world_rank, ping_pong_count, partner_rank);
  fflush(stdout); // 强制刷新输出缓冲
} else {
  MPI_Recv(&ping_pong_count, 1, MPI_INT, partner_rank, 0, MPI_COMM_WORLD,
           MPI_STATUS_IGNORE);
  printf("%d received ping_pong_count %d from %d\n",
         world_rank, ping_pong_count, partner_rank);
  fflush(stdout); // 强制刷新输出缓冲
}

额外验证方案

若修改后仍有问题,可检查以下配置:

  • 执行命令时添加--oversubscribe参数:mpirun -np 2 --oversubscribe ./ping_pong(适用于核心数不足的场景);
  • 检查OpenMPI通信驱动,确保使用本地高效通信模块(如shared或tcp)。

内容的提问来源于stack exchange,提问作者Lindberg Young

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 03:13:10