链接时突发maxrregcount警告与未定义引用错误的排查求助
CUDA分离编译dlink阶段警告及未定义引用问题排查与解决
我维护着一个C++风格的CUDA API封装库,当前版本已经过充分测试,配套示例程序也有不少用户在使用。但最近怪事发生了——明明没提交任何新代码,编译示例程序的dlink阶段突然冒出一堆NVCC警告,后续链接还触发了未定义引用错误,折腾了我好一阵子。
先看看遇到的具体问题
1. dlink阶段的重复Beta特性警告
dlink阶段能完成,但反复输出相同的ptxas信息:
/path/to/nvcc /path/to/cuda-api-wrappers/examples/modified_cuda_samples/vectorAdd/vectorAdd.cu -dc -o /path/to/cuda-api-wrappers/CMakeFiles/vectorAdd.dir/examples/modified_cuda_samples/vectorAdd/./vectorAdd_generated_vectorAdd.cu.o -ccbin /opt/gcc-5.4.0/bin/gcc -m64 -gencode arch=compute_52,code=compute_52 --std=c++11 -Xcompiler -Wall -O3 -DNDEBUG -DNVCC -I/path/to/cuda/include -I/path/to/cuda-api-wrappers/src /path/to/nvcc -gencode arch=compute_52,code=compute_52 --std=c++11 -Xcompiler -Wall -O3 -DNDEBUG -m64 -ccbin /opt/gcc-5.4.0/bin/gcc -dlink /export/path/to/cuda-api-wrappers/CMakeFiles/vectorAdd.dir/examples/modified_cuda_samples/vectorAdd/./vectorAdd_generated_vectorAdd.cu.o /path/to/cuda/lib64/libcublas_device.a -o /export/path/to/cuda-api-wrappers/CMakeFiles/vectorAdd.dir/./vectorAdd_intermediate_link.o @O@ptxas info : 'device-function-maxrregcount' is a BETA feature @O@ptxas info : 'device-function-maxrregcount' is a BETA feature @O@ptxas info : 'device-function-maxrregcount' is a BETA feature ... this repeats many times ...
2. 链接阶段的未定义引用错误
dlink过了,但最终链接时抛出一堆类似这样的错误:
/opt/gcc-5.4.0/bin/g++ -Wall -Wpedantic -O2 -DNDEBUG -L/path/to/cuda/lib64 -rdynamic CMakeFiles/vectorAdd.dir/examples/modified_cuda_samples/vectorAdd/vectorAdd_generated_vectorAdd.cu.o CMakeFiles/vectorAdd.dir/vectorAdd_intermediate_link.o -o examples/bin/vectorAdd lib/libcuda-api-wrappers.a -Wl,-Bstatic -lcudart_static -Wl,-Bdynamic -lpthread -ldl -lrt -lnvToolsExt -Wl,-Bstatic -lcudadevrt -Wl,-Bdynamic CMakeFiles/vectorAdd.dir/vectorAdd_intermediate_link.o: In function `__cudaRegisterLinkedBinary_25_cublas_compute_70_cpp1_ii_f0559976': link.stub:(.text+0xe0): undefined reference to `__fatbinwrap_25_cublas_compute_70_cpp1_ii_f0559976' CMakeFiles/vectorAdd.dir/vectorAdd_intermediate_link.o: In function `__cudaRegisterLinkedBinary_25_xerbla_compute_70_cpp1_ii_cd7f3ad3': link.stub:(.text+0x190): undefined reference to `__fatbinwrap_25_xerbla_compute_70_cpp1_ii_cd7f3ad3' CMakeFiles/vectorAdd.dir/vectorAdd_intermediate_link.o: In function `__cudaRegisterLinkedBinary_23_nrm2_compute_70_cpp1_ii_8edbce95': link.stub:(.text+0x240): undefined reference to `__fatbinwrap_23_nrm2_compute_70_cpp1_ii_8edbce95' ... more undefined reference errors here ...
问题成因分析
结合我的环境(CUDA 9.1、SM 5.2设备,无SM 7.0硬件),仔细看错误信息后发现问题根源:
- Beta警告的来源:dlink命令里链接了
libcublas_device.a——这个库是给设备端调用cuBLAS用的,而CUDA 9.x里这个库的代码使用了device-function-maxrregcount这个Beta特性,所以ptxas会反复输出警告,哪怕我自己的代码根本没碰这个特性。 - 未定义引用的原因:我编译时只指定了SM 5.2的目标(
-gencode arch=compute_52,code=compute_52),但libcublas_device.a里包含了针对SM 7.0的fatbin代码,这些代码对应的符号在我的编译产物里找不到实现,就出现了__fatbinwrap_*的未定义引用。
更关键的是:我的vectorAdd示例根本不需要设备端cuBLAS,完全是多余链接了这个库导致的问题!
解决方法
- 移除不必要的
libcublas_device.a链接:检查你的CMake配置,找到哪里引入了libcublas_device.a,直接删掉它。对于vectorAdd这种基础示例,完全不需要链接设备端cuBLAS库。 - 清理缓存重新构建:修改CMake配置后,一定要清理CMakeCache.txt,然后重新生成构建文件,确保修改生效:
rm CMakeCache.txt cmake .. make clean && make - 若确实需要设备端cuBLAS的兼容处理(可选):如果你的项目真的需要设备端cuBLAS,那得在编译时加入SM 7.0的目标代码生成,即使你没有SM 7.0设备,CUDA也能生成兼容代码。修改
-gencode参数,加上:
不过这会增加二进制大小,对于不需要的场景不建议这么做。-gencode arch=compute_70,code=sm_70
总结
这次的问题完全是误链接了不必要的设备端cuBLAS库导致的,和我的代码本身无关。排查时一定要仔细看链接命令里的库列表,尤其是那些带_device后缀的CUDA库,它们是针对设备端调用的,普通的主机端CUDA程序根本不需要碰。
内容的提问来源于stack exchange,提问作者einpoklum
相关产品推荐
相关产品推荐

