为何#pragma omp loop for会引发编译错误?
OpenMP与OpenACC混合编程编译错误排查
编译错误信息
编译命令及错误输出:
$ mpic++ -mcmodel=medium -fopenmp -acc -ta=tesla:managed -Minfo=accel mp_acc.c -o mp_acc "mp_acc.c", line 49: error: invalid text in pragma #pragma omp loop for ^ 1 error detected in the compilation of "mp_acc.c". $
编译器版本
$ mpic++ --version nvc++ 22.1-0 64-bit target on x86-64 Linux -tp skylake-avx512 NVIDIA Compilers and Tools Copyright (c) 2022, NVIDIA CORPORATION & AFFILIATES. All rights reserved.
完整测试代码
#include <iostream> #include <stdio.h> #include <stdlib.h> #include <cstdlib> #include <string> #include <mpi.h> #include <omp.h> #include <openacc.h> #include <bits/stdc++.h> #include <sys/stat.h> using namespace std; void* allocMatrix (int nRow, int nCol) { void* restrict m = malloc (sizeof(int[nRow][nCol])); return(m); } #pragma acc routine gang void* func(void* a, int nrows, int ncols) { return(a); } int main(int argc, char *argv[]) { int nrows = 5; int ncols = 3; int (*a)[ncols] = (int (*)[ncols])allocMatrix(nrows, ncols); int* restrict ta = (int*)malloc(nrows * sizeof(int)); for ( int i=0; i<nrows; i++ ) { for ( int j=0; j<ncols; j++ ) { a[i][j] = 1; } } for ( int i=0; i<nrows; i++ ) { for ( int j=0; j<ncols; j++ ) { cout << a[i][j] << " "; } cout << endl; } #pragma omp parallel num_threads() { size_t tid = omp_get_thread_num(); #pragma omp loop for (int i = 0; i < nrows; ++i) { #pragma acc parallel deviceptr(a,nrows,ncols) async(tid) { a = (int (*)[ncols]) func(a, nrows, ncols); } } #pragma acc wait } for ( int i=0; i<nrows; i++ ) { for ( int j=0; j<ncols; j++ ) { cout << a[i][j] << " "; } cout << endl; } memset( a, 0, nrows*ncols*sizeof(int) ); for ( int i=0; i<nrows; i++ ) { for ( int j=0; j<ncols; j++ ) { cout << a[i][j] << " "; } cout << endl; } free(a); }
错误分析与修复方案
1. 核心编译错误:OpenMP指令语法错误
错误提示的#pragma omp loop for是非法语法。OpenMP中,loop是for指令的别名(OpenMP 4.0及以后支持),二者不能同时使用,二选一即可:
- 写成
#pragma omp for(通用写法) - 或者
#pragma omp loop
2. 其他潜在问题及修复
(1) OpenMP并行区域线程数未指定
#pragma omp parallel num_threads()中num_threads缺少参数,需指定线程数量,示例:
// 指定固定线程数 #pragma omp parallel num_threads(4) // 或使用系统默认最大线程数 #pragma omp parallel num_threads(omp_get_max_threads())
(2) OpenACC deviceptr使用错误
#pragma acc parallel deviceptr(a,nrows,ncols)中,nrows和ncols是标量变量,不需要加入deviceptr列表(deviceptr仅用于标记设备端指针变量),正确写法:
#pragma acc parallel deviceptr(a) async((int)tid)
(3) OpenACC async参数类型不匹配
async(tid)中tid是size_t类型,而OpenACC要求async参数为int,需强制转换:
async((int)tid)
(4) MPI未初始化与清理
代码包含MPI头文件但未调用初始化/清理函数,会导致运行时错误,需在main开头和结尾添加:
MPI_Init(&argc, &argv); // ... 原有代码 ... MPI_Finalize();
(5) C++变长数组兼容性问题
allocMatrix中sizeof(int[nRow][nCol])属于C语言变长数组(VLA)特性,并非C++标准语法,仅部分编译器支持扩展。更兼容的写法:
void* allocMatrix (int nRow, int nCol) { void* restrict m = malloc(nRow * nCol * sizeof(int)); return(m); }
修复后的核心代码片段
#pragma omp parallel num_threads(omp_get_max_threads()) { int tid = (int)omp_get_thread_num(); #pragma omp for for (int i = 0; i < nrows; ++i) { #pragma acc parallel deviceptr(a) async(tid) { a = (int (*)[ncols]) func(a, nrows, ncols); } } #pragma acc wait }
内容的提问来源于stack exchange,提问作者Mark Bower
相关产品推荐
相关产品推荐

