OpenMP GPU卸载加速时C语言调用math.h函数的正确方案
问题
在基于OpenMP实现CPU/GPU可选加速的C语言代码中,矩阵乘法的GPU卸载运行正常,但调用math.h中expf实现的sigmoid函数时,出现运行时错误:libgomp: pointer target not mapped for attach。已通过CMake配置GCC的NVIDIA GPU卸载参数,尝试过-ffast-math但无效,需要解决OpenMP GPU卸载场景下正确使用math.h函数的问题。
相关代码示例
矩阵乘法函数
/* ... */ int numeric_matmul(const float_t *pt_a, const float_t *pt_b, float_t *pt_c, uintmax_t t_m, uintmax_t t_k, uintmax_t t_n) { #ifdef _OPENMP #pragma omp target teams distribute parallel for collapse(2) schedule(dynamic) map(to: pt_a[0 : t_m * t_k], pt_b[0 : t_k * t_n]) map(from: pt_c[0 : t_m * t_n]) #endif for(uintmax_t l_i = 0; l_i < t_m; l_i++) { for(uintmax_t l_j = 0; l_j < t_n; l_j++) { /* Compute the sum. */ float_t l_sum = 0.0; for(uintmax_t l_p = 0; l_p < t_k; l_p++) l_sum += pt_a[l_i * t_k + l_p] * pt_b[l_p * t_n + l_j]; /* Store the result. */ pt_c[l_i * t_n + l_j] = l_sum; } } /* Return with success. */ return 0; }
sigmoid函数
/** * @brief Perform the sigmoid function on a value. * @param t_x The input value. * @param pt_y The output value. * @return The result status code. In this case, it'll always return 0. */ static inline int numeric_sigmoid(float_t t_x, float_t *pt_y) { /* Set the output value to the sigmoid of the input value. */ *pt_y = 1.0 / (1.0 + expf(-t_x)); /* Return with success. */ return 0; }
GPU卸载调用代码
#pragma omp target teams distribute parallel for schedule(dynamic) map(to: pt_feedforward->ppt_hidden_layer_bias_buffer[l_i][0 : l_next_layer_activation_buffer_size]) map(from: pl_next_layer_activation_buffer[0 : l_next_layer_activation_buffer_size]) for(uintmax_t l_j = 0; l_j < l_next_layer_activation_buffer_size; l_j++) { pl_next_layer_activation_buffer[l_j] += pt_feedforward->ppt_hidden_layer_bias_buffer[l_i][l_j]; numeric_sigmoid(pl_next_layer_activation_buffer[l_j], &pl_next_layer_activation_buffer[l_j]); }
编译参数
cmake -DCMAKE_C_COMPILER=gcc -DCMAKE_C_FLAGS="-fopenmp -foffload=nvptx-none -foffload-options=-misa=sm_80 -fcf-protection=none -fno-stack-protector -no-pie" ..
解决方案
核心原因
GCC的OpenMP GPU卸载不会自动为math.h中的主机端函数生成设备端实现,直接调用会导致设备端找不到对应函数,进而触发指针映射错误。另外,numeric_sigmoid作为static inline函数,需要明确标记让编译器为设备端生成代码。
具体修复步骤
标记函数可在设备端执行
使用OpenMP的#pragma omp declare target指令,把sigmoid函数和它依赖的expf标记为设备端可用,两种方式任选:- 直接把
sigmoid函数包裹在声明块中:#pragma omp declare target static inline int numeric_sigmoid(float_t t_x, float_t *pt_y) { *pt_y = 1.0 / (1.0 + expf(-t_x)); return 0; } #pragma omp end declare target - 单独声明
expf设备端可用:#pragma omp declare target(expf)
- 直接把
确保内联逻辑在设备端生效
添加declare target后,编译器会自动处理static inline函数的设备端内联,避免函数调用时的指针映射问题。优化编译参数(可选)
把-ffast-math加入编译参数,配合declare target能进一步优化设备端数学函数的性能:cmake -DCMAKE_C_COMPILER=gcc -DCMAKE_C_FLAGS="-fopenmp -foffload=nvptx-none -foffload-options=-misa=sm_80 -fcf-protection=none -fno-stack-protector -no-pie -ffast-math" ..
内容的提问来源于Stack Exchange,提问作者Matthew G.
相关产品推荐
相关产品推荐

