nvcc使用__gnu_parallel::sort()报cmath重定义错误如何解决
问题背景
当前使用环境如下:
- 操作系统:Ubuntu 16.04
- 编译工具链:NVCC 7.5、GCC 5.4.0
待编译的testFile.cu代码如下:
#include <math.h> #include <parallel/algorithm> void main(){ <some work is beeing donbe here> __gnu_parallel::sort(vector.begin(),vector.end(),<comparator function>); }
编译时出现如下重定义报错:
/usr/include/c++/5/tr1/cmath(424): error: function "acosh(float)" has already been defined /usr/include/c++/5/tr1/cmath(442): error: function "asinh(float)" has already been defined /usr/include/c++/5/tr1/cmath(461): error: function "atanh(float)" has already been defined /usr/include/c++/5/tr1/cmath(477): error: function "cbrt(float)" has already been defined /usr/include/c++/5/tr1/cmath(495): error: function "copysign(float, float)" has already been defined /usr/include/c++/5/tr1/cmath(516): error: function "erf(float)" has already been defined /usr/include/c++/5/tr1/cmath(532): error: function "erfc(float)" has already been defined /usr/include/c++/5/tr1/cmath(550): error: function "exp2(float)" has already been defined /usr/include/c++/5/tr1/cmath(566): error: function "expm1(float)" has already been defined /usr/include/c++/5/tr1/cmath(608): error: function "fdim(float, float)" has already been defined
目标是在NVCC编译环境下正常调用__gnu_parallel::sort()。
报错原因
NVCC 7.5对GCC 5的TR1数学库头文件兼容性存在缺陷:<parallel/algorithm>会引入<tr1/cmath>头文件,和提前引入的<math.h>中声明的C99浮点数学函数产生重定义冲突。
解决方案
有两种可行方案,优先选择第二种,稳定性更高:
方案1:快速修复(单文件编译可用)
调整头文件包含顺序,将<parallel/algorithm>放在所有标准库头文件的最前面,同时编译时添加宏定义关闭GCC TR1数学库的C99函数声明,避免重定义。
调整后的代码头文件部分:// 优先包含parallel/algorithm #include <parallel/algorithm> #include <math.h> // 其余头文件 int main(){ // 注意标准C++中main函数返回值类型应为int,不要写void main // 业务逻辑 __gnu_parallel::sort(vector.begin(),vector.end(),<comparator function>); return 0; }对应编译命令:
nvcc testFile.cu -o test -D_GLIBCXX_USE_C99_MATH_TR1=0 -Xcompiler -fopenmp -O3注意
__gnu_parallel::sort依赖OpenMP,必须加-fopenmp编译选项才能启用多线程。方案2:分离编译(推荐,无兼容性隐患)
NVCC本身对GNU专属的并行扩展头文件解析支持有限,最稳妥的方式是把CPU端的并行排序逻辑和CUDA设备端代码拆分,分别用GCC和NVCC编译后再链接,彻底规避头文件解析冲突。- 新建
sort_wrapper.cpp存放CPU端并行排序逻辑:
#include <vector> #include <parallel/algorithm> // 此处替换为你实际使用的元素类型和比较器类型 void parallel_sort(std::vector<YourDataType>& vec, YourComparator cmp) { __gnu_parallel::sort(vec.begin(), vec.end(), cmp); }- 修改
testFile.cu,移除<parallel/algorithm>头文件,声明外部排序函数:
#include <math.h> #include <vector> // 声明外部实现的并行排序函数 void parallel_sort(std::vector<YourDataType>& vec, YourComparator cmp); int main(){ // 其余业务逻辑、CUDA核调用逻辑 parallel_sort(your_vector, your_comparator); return 0; }- 分两步编译后链接:
# 用GCC编译CPU端并行排序代码 g++ -c sort_wrapper.cpp -o sort_wrapper.o -O3 -fopenmp # 用NVCC编译CUDA代码,链接CPU端目标文件 nvcc testFile.cu sort_wrapper.o -o test -O3 -Xcompiler -fopenmp- 新建
内容的提问来源于stack exchange,提问作者Ankur Ghoshal
相关产品推荐
相关产品推荐

