You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

nvcc使用__gnu_parallel::sort()报cmath重定义错误如何解决

问题背景

当前使用环境如下:

  • 操作系统:Ubuntu 16.04
  • 编译工具链:NVCC 7.5、GCC 5.4.0

待编译的testFile.cu代码如下:

#include <math.h>
#include <parallel/algorithm>
void main(){
    
      <some work is beeing donbe here>

      __gnu_parallel::sort(vector.begin(),vector.end(),<comparator function>);
}

编译时出现如下重定义报错:

/usr/include/c++/5/tr1/cmath(424): error: function "acosh(float)" has already been defined
/usr/include/c++/5/tr1/cmath(442): error: function "asinh(float)" has already been defined
/usr/include/c++/5/tr1/cmath(461): error: function "atanh(float)" has already been defined
/usr/include/c++/5/tr1/cmath(477): error: function "cbrt(float)" has already been defined
/usr/include/c++/5/tr1/cmath(495): error: function "copysign(float, float)" has already been defined
/usr/include/c++/5/tr1/cmath(516): error: function "erf(float)" has already been defined
/usr/include/c++/5/tr1/cmath(532): error: function "erfc(float)" has already been defined
/usr/include/c++/5/tr1/cmath(550): error: function "exp2(float)" has already been defined
/usr/include/c++/5/tr1/cmath(566): error: function "expm1(float)" has already been defined
/usr/include/c++/5/tr1/cmath(608): error: function "fdim(float, float)" has already been defined

目标是在NVCC编译环境下正常调用__gnu_parallel::sort()。

报错原因

NVCC 7.5对GCC 5的TR1数学库头文件兼容性存在缺陷:<parallel/algorithm>会引入<tr1/cmath>头文件,和提前引入的<math.h>中声明的C99浮点数学函数产生重定义冲突。

解决方案

有两种可行方案,优先选择第二种,稳定性更高:

  • 方案1:快速修复(单文件编译可用)
    调整头文件包含顺序,将<parallel/algorithm>放在所有标准库头文件的最前面,同时编译时添加宏定义关闭GCC TR1数学库的C99函数声明,避免重定义。
    调整后的代码头文件部分:

    // 优先包含parallel/algorithm
    #include <parallel/algorithm>
    #include <math.h>
    // 其余头文件
    int main(){ // 注意标准C++中main函数返回值类型应为int,不要写void main
        // 业务逻辑
        __gnu_parallel::sort(vector.begin(),vector.end(),<comparator function>);
        return 0;
    }
    

    对应编译命令:

    nvcc testFile.cu -o test -D_GLIBCXX_USE_C99_MATH_TR1=0 -Xcompiler -fopenmp -O3
    

    注意__gnu_parallel::sort依赖OpenMP,必须加-fopenmp编译选项才能启用多线程。

  • 方案2:分离编译(推荐,无兼容性隐患)
    NVCC本身对GNU专属的并行扩展头文件解析支持有限,最稳妥的方式是把CPU端的并行排序逻辑和CUDA设备端代码拆分,分别用GCC和NVCC编译后再链接,彻底规避头文件解析冲突。

    1. 新建sort_wrapper.cpp存放CPU端并行排序逻辑:
    #include <vector>
    #include <parallel/algorithm>
    // 此处替换为你实际使用的元素类型和比较器类型
    void parallel_sort(std::vector<YourDataType>& vec, YourComparator cmp) {
        __gnu_parallel::sort(vec.begin(), vec.end(), cmp);
    }
    
    1. 修改testFile.cu,移除<parallel/algorithm>头文件,声明外部排序函数:
    #include <math.h>
    #include <vector>
    // 声明外部实现的并行排序函数
    void parallel_sort(std::vector<YourDataType>& vec, YourComparator cmp);
    
    int main(){
        // 其余业务逻辑、CUDA核调用逻辑
        parallel_sort(your_vector, your_comparator);
        return 0;
    }
    
    1. 分两步编译后链接:
    # 用GCC编译CPU端并行排序代码
    g++ -c sort_wrapper.cpp -o sort_wrapper.o -O3 -fopenmp
    # 用NVCC编译CUDA代码,链接CPU端目标文件
    nvcc testFile.cu sort_wrapper.o -o test -O3 -Xcompiler -fopenmp
    

内容的提问来源于stack exchange,提问作者Ankur Ghoshal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 14:51:17