如何避免Thread Sanitizer对扩展Lambda实现的误报?
解决GCC Thread Sanitizer中NVCC扩展Lambda的误报问题
我尝试用GCC的Thread Sanitizer(-fsanitize=thread)检查应用的数据竞争,但输出被大量NVCC扩展Lambda实现导致的误报淹没。
示例代码
#include <thread> #include <array> int main() { std::array<std::thread, 2> threads; for (unsigned i = 0; i < 2; ++i) threads[i] = std::thread{[] { [] __host__ __device__ {}; }}; threads[0].join(); threads[1].join(); return 0; }
Thread Sanitizer警告输出
WARNING: ThreadSanitizer: data race (pid=1) Write of size 8 at 0x0000004a45e0 by thread T2: #0 __nv_hdl_wrapper_t<main()::<lambda()>::<lambda()> > /app/nvcc_internal_extended_lambda_implementation:296 (output.s+0x405137) #1 __nv_hdl_create_wrapper<main()::<lambda()>::<lambda()> > /app/nvcc_internal_extended_lambda_implementation:347 (output.s+0x405041) #2 operator() /app/example.cu:8 (output.s+0x404d52) #3 __invoke_impl<void, main()::<lambda()> > /opt/compiler-explorer/gcc-10.2.0/include/c++/10.2.0/bits/invoke.h:60 (output.s+0x405584) #4 __invoke<main()::<lambda()> > /opt/compiler-explorer/gcc-10.2.0/include/c++/10.2.0/bits/invoke.h:95 (output.s+0x4054f1) #5 _M_invoke<0> /opt/compiler-explorer/gcc-10.2.0/include/c++/10.2.0/thread:264 (output.s+0x405456) #6 operator() /opt/compiler-explorer/gcc-10.2.0/include/c++/10.2.0/thread:271 (output.s+0x405400) #7 _M_run /opt/compiler-explorer/gcc-10.2.0/include/c++/10.2.0/thread:215 (output.s+0x4053ba) #8 <null> <null> (libstdc++.so.6+0xd6de3) Previous write of size 8 at 0x0000004a45e0 by thread T1: #0 __nv_hdl_wrapper_t<main()::<lambda()>::<lambda()> > /app/nvcc_internal_extended_lambda_implementation:296 (output.s+0x405137) #1 __nv_hdl_create_wrapper<main()::<lambda()>::<lambda()> > /app/nvcc_internal_extended_lambda_implementation:347 (output.s+0x405041) #2 operator() /app/example.cu:8 (output.s+0x404d52) #3 __invoke_impl<void, main()::<lambda()> > /opt/compiler-explorer/gcc-10.2.0/include/c++/10.2.0/bits/invoke.h:60 (output.s+0x405584) #4 __invoke<main()::<lambda()> > /opt/compiler-explorer/gcc-10.2.0/include/c++/10.2.0/bits/invoke.h:95 (output.s+0x4054f1) #5 _M_invoke<0> /opt/compiler-explorer/gcc-10.2.0/include/c++/10.2.0/thread:264 (output.s+0x405456) #6 operator() /opt/compiler-explorer/gcc-10.2.0/include/c++/10.2.0/thread:271 (output.s+0x405400) #7 _M_run /opt/compiler-explorer/gcc-10.2.0/include/c++/10.2.0/thread:215 (output.s+0x4053ba) #8 <null> <null> (libstdc++.so.6+0xd6de3) Location is global '(anonymous namespace)::__nv_hdl_helper<__nv_dl_tag<int (*)(), &main, 1u>, void>::fp_noobject_caller' of size 8 at 0x0000004a45e0 (output.s+0x0000004a45e0) Thread T2 (tid=4, running) created by main thread at: #0 pthread_create ../../../../src/libsanitizer/tsan/tsan_interceptors_posix.cpp:962 (libtsan.so.0+0x5ea79) #1 std::thread::_M_start_thread(std::unique_ptr<std::thread::_State, std::default_delete<std::thread::_State> >, void (*)()) <null> (libstdc++.so.6+0xd70a8) #2 main /app/example.cu:8 (output.s+0x404db3) Thread T1 (tid=3, finished) created by main thread at: #0 pthread_create ../../../../src/libsanitizer/tsan/tsan_interceptors_posix.cpp:962 (libtsan.so.0+0x5ea79) #1 std::thread::_M_start_thread(std::unique_ptr<std::thread::_State, std::default_delete<std::thread::_State> >, void (*)()) <null> (libstdc++.so.6+0xd70a8) #2 main /app/example.cu:8 (output.s+0x404db3) SUMMARY: ThreadSanitizer: data race /app/nvcc_internal_extended_lambda_implementation:296 in __nv_hdl_wrapper_t<main()::<lambda()>::<lambda()> >
问题根源
误报来自NVCC扩展Lambda实现中,创建Lambda时对静态变量的无锁赋值。核心代码片段如下:
template <typename Tag, typename OpFuncR, typename ...OpFuncArgs> struct __nv_hdl_helper { typedef void * (*fp_copier_t)(void *); typedef OpFuncR (*fp_caller_t)(void *, OpFuncArgs...); typedef void (*fp_deleter_t) (void *); typedef OpFuncR (*fp_noobject_caller_t)(OpFuncArgs...); static fp_copier_t fp_copier; static fp_deleter_t fp_deleter; static fp_caller_t fp_caller; static fp_noobject_caller_t fp_noobject_caller; }; /* .... */ typedef OpFuncR(__opfunc_t)(OpFuncArgs...); template <typename Lambda> __nv_hdl_wrapper_t(Tag, Lambda &&lam, F1 in1 , F2 in2 ) : f1(in1) ,f2(in2) { __nv_hdl_helper<Tag, OpFuncR, OpFuncArgs...>::fp_noobject_caller = lam; }
虽然多线程写入的是相同值(实际无影响,但属于未定义行为),但TSAN会将其标记为数据竞争,且我们无法修改NVCC的这部分实现。
解决方案
1. 使用TSAN抑制文件
创建一个抑制文件(例如tsan_suppressions.txt),添加针对NVCC扩展Lambda相关函数的规则:
race:__nv_hdl_wrapper_t race:__nv_hdl_create_wrapper race:__nv_hdl_helper
编译运行时通过-fsanitize=thread -fsanitize-blacklist=tsan_suppressions.txt(GCC 10+可用)或TSAN_OPTIONS="suppressions=tsan_suppressions.txt"环境变量加载抑制规则。
2. 手动标记无竞争访问(如果可行)
如果你能在代码中包裹Lambda创建逻辑,可以使用TSAN的__tsan_acquire和__tsan_release宏标记该区域为安全:
#include <sanitizer/tsan_interface.h> // ... threads[i] = std::thread{[] { __tsan_acquire(nullptr); [] __host__ __device__ {}; __tsan_release(nullptr); }}; // ...
这种方式通过告诉TSAN该区域的访问是同步的,避免误报,但需要修改业务代码。
3. 过滤TSAN输出
运行程序时,通过管道过滤掉特定模式的误报:
./your_program 2>&1 | grep -v "__nv_hdl_"
这是最简单的临时方案,但可能会误过滤真正的问题,适合快速排查。
内容的提问来源于Stack Exchange,提问作者Lukas Lang
相关产品推荐
相关产品推荐

