You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编译含递归函数的OpenMP target GPU内核?

问题:GCC编译含递归函数的OpenMP target内核失败

使用OpenMP target构造在GPU设备上运行循环时,GCC编译器无法编译包含递归函数的代码,取消递归调用后编译恢复正常。以下是复现问题的详细信息及解决方案:

示例代码

int func(int par){
 if(par<0) return par;
 return func(par--);
}

int main(void)
{
  int N = 100;
      #pragma omp target teams distribute parallel for 
      for (int i = 0; i<N; i++) {
          func(i);
  }
}

编译命令

gcc-12  -foffload=nvptx-none -fcf-protection=none -fno-stack-protector -foffload=-misa=sm_35  -fopenmp test.c -o testcpp.exe

错误信息

test.c:5:5: error: alias definitions not supported in this configuration
    5 | int func(int par){
      |     ^
mkoffload: fatal error: x86_64-linux-gnu-accel-nvptx-none-gcc-12 returned 1 exit status
compilation terminated.
lto-wrapper: fatal error: /usr/lib/gcc/x86_64-linux-gnu/12//accel/nvptx-none/mkoffload returned 1 exit status
compilation terminated.
/usr/bin/ld: error: lto-wrapper failed
collect2: error: ld returned 1 exit status

解决方案

  • 标记递归函数为设备函数:
    给递归函数添加#pragma omp declare target指令,明确告知编译器该函数需要被编译到GPU设备端。同时修正原代码中par--的错误(后置递减会导致无限递归),修改后的代码如下:

    #pragma omp declare target
    int func(int par){
     if(par<0) return par;
     return func(par - 1);
    }
    #pragma omp end declare target
    
  • 关闭链接时优化(LTO):
    GCC在启用LTO时处理设备端递归函数存在兼容性问题,编译时添加-flto=off参数关闭LTO,调整后的编译命令:

    gcc-12  -foffload=nvptx-none -fcf-protection=none -fno-stack-protector -foffload=-misa=sm_35  -fopenmp -flto=off test.c -o testcpp.exe
    
  • 升级GCC版本:
    GCC 12对OpenMP target递归函数的支持存在局限,升级到GCC 13及以上版本可获得更好的设备端递归兼容性。

内容的提问来源于stack exchange,提问作者Antonio Ragagnin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 16:23:27