You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenMP GPU卸载加速时C语言调用math.h函数的正确方案

问题

在基于OpenMP实现CPU/GPU可选加速的C语言代码中,矩阵乘法的GPU卸载运行正常,但调用math.h中expf实现的sigmoid函数时,出现运行时错误:libgomp: pointer target not mapped for attach。已通过CMake配置GCC的NVIDIA GPU卸载参数,尝试过-ffast-math但无效,需要解决OpenMP GPU卸载场景下正确使用math.h函数的问题。

相关代码示例

矩阵乘法函数

/* ... */
int numeric_matmul(const float_t *pt_a, const float_t *pt_b, float_t *pt_c, uintmax_t t_m, uintmax_t t_k, uintmax_t t_n)
{
#ifdef _OPENMP
#pragma omp target teams distribute parallel for collapse(2) schedule(dynamic) map(to: pt_a[0 : t_m * t_k], pt_b[0 : t_k * t_n]) map(from: pt_c[0 : t_m * t_n])
#endif
    for(uintmax_t l_i = 0; l_i < t_m; l_i++)
    {
        for(uintmax_t l_j = 0; l_j < t_n; l_j++)
        {
/* Compute the sum. */
            float_t l_sum = 0.0;
            for(uintmax_t l_p = 0; l_p < t_k; l_p++) l_sum += pt_a[l_i * t_k + l_p] * pt_b[l_p * t_n + l_j];

/* Store the result. */
            pt_c[l_i * t_n + l_j] = l_sum;
        }
    }

/* Return with success. */
    return 0;
}

sigmoid函数

/**
 *  @brief Perform the sigmoid function on a value.
 *  @param t_x The input value.
 *  @param pt_y The output value.
 *  @return The result status code. In this case, it'll always return 0.
 */
static inline int numeric_sigmoid(float_t t_x, float_t *pt_y)
{
/* Set the output value to the sigmoid of the input value. */
    *pt_y = 1.0 / (1.0 + expf(-t_x));

/* Return with success. */
    return 0;
}

GPU卸载调用代码

#pragma omp target teams distribute parallel for schedule(dynamic) map(to: pt_feedforward->ppt_hidden_layer_bias_buffer[l_i][0 : l_next_layer_activation_buffer_size]) map(from: pl_next_layer_activation_buffer[0 : l_next_layer_activation_buffer_size])
for(uintmax_t l_j = 0; l_j < l_next_layer_activation_buffer_size; l_j++)
{
    pl_next_layer_activation_buffer[l_j] += pt_feedforward->ppt_hidden_layer_bias_buffer[l_i][l_j];
    numeric_sigmoid(pl_next_layer_activation_buffer[l_j], &pl_next_layer_activation_buffer[l_j]);
}

编译参数

cmake -DCMAKE_C_COMPILER=gcc -DCMAKE_C_FLAGS="-fopenmp -foffload=nvptx-none -foffload-options=-misa=sm_80 -fcf-protection=none -fno-stack-protector -no-pie" ..
解决方案

核心原因

GCC的OpenMP GPU卸载不会自动为math.h中的主机端函数生成设备端实现,直接调用会导致设备端找不到对应函数,进而触发指针映射错误。另外,numeric_sigmoid作为static inline函数,需要明确标记让编译器为设备端生成代码。

具体修复步骤

  1. 标记函数可在设备端执行
    使用OpenMP的#pragma omp declare target指令,把sigmoid函数和它依赖的expf标记为设备端可用,两种方式任选:

    • 直接把sigmoid函数包裹在声明块中:
      #pragma omp declare target
      static inline int numeric_sigmoid(float_t t_x, float_t *pt_y)
      {
          *pt_y = 1.0 / (1.0 + expf(-t_x));
          return 0;
      }
      #pragma omp end declare target
      
    • 单独声明expf设备端可用:
      #pragma omp declare target(expf)
      
  2. 确保内联逻辑在设备端生效
    添加declare target后,编译器会自动处理static inline函数的设备端内联,避免函数调用时的指针映射问题。

  3. 优化编译参数(可选)
    把-ffast-math加入编译参数,配合declare target能进一步优化设备端数学函数的性能:

    cmake -DCMAKE_C_COMPILER=gcc -DCMAKE_C_FLAGS="-fopenmp -foffload=nvptx-none -foffload-options=-misa=sm_80 -fcf-protection=none -fno-stack-protector -no-pie -ffast-math" ..
    

内容的提问来源于Stack Exchange,提问作者Matthew G.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 14:25:15