如何忽略Valgrind DRD中std::future的误报数据竞争?
std::promise/std::future被Valgrind DRD误报数据竞争的原因与解决办法
问题背景
我正在编写库代码,用std::future将lambda(或std::function/可调用对象)的执行上下文转移到另一个线程,以此避免数据竞争——因为该std::function的操作几乎完全与另一线程管理的数据相关。这本质上是带显式异常传递的std::promise/std::future标准用法,代码结构如下。
但使用Valgrind工具集的DRD测试时,出现了奇怪的数据竞争报告。根据cppreference文档,这些操作本应是线程安全的,无需额外同步。为测试起见,我甚至在std::promise/future的访问周围添加了std::mutex和lock_guard,结果更令人困惑——我自己添加的mutex反而被报告为冲突操作的原因。GCC 12和Clang 13均测试过,结果一致。
核心问题
- 为什么Valgrind DRD会标记
std::promise/std::future的操作不安全? - 如何解决这个误报问题?
- 有没有办法将这些合法操作加入白名单、证明其合法性或隐藏误报?
环境:Debian系统,libstdc++6 12.1.0-5
Valgrind DRD报错日志
==45452== Thread 2: ==45452== Conflicting load by thread 2 at 0x05c64800 size 4 ==45452== at 0x4928B55: load (atomic_base.h:488) ==45452== by 0x4928B55: _M_load (atomic_futex.h:86) ==45452== by 0x4928B55: std::__atomic_futex_unsigned<2147483648u>::_M_load_and_test_until(unsigned int, unsigned int, bool, std::memory_order, bool, std::chrono::duration<long, std::ratio<1l, 1l> >, std::chrono::duration<long, std::ratio<1l, 1000000000l> >) (atomic_futex.h:113) ==45452== by 0x49287FC: std::__atomic_futex_unsigned<2147483648u>::_M_load_and_test(unsigned int, unsigned int, bool, std::memory_order) (atomic_futex.h:158) ==45452== by 0x4928697: _M_load_when_equal (atomic_futex.h:212) ==45452== by 0x4928697: std::__future_base::_State_baseV2::wait() (future:337) ==45452== by 0x4959849: std::__basic_future<unsigned long>::_M_get_result() const (future:720) ==45452== by 0x4955E55: std::future<unsigned long>::get() (future:806) ==45452== by 0x4954AD5: RunOnIoThread(std::function<unsigned long ()>, unsigned long) (schedlib.cc:335) ... ==45452== Address 0x5c64800 is at offset 32 from 0x5c647e0. Allocation context: ==45452== at 0x484514F: operator new(unsigned long) (in /usr/libexec/valgrind/vgpreload_drd-amd64-linux.so) ==45452== by 0x4925A4C: std::__new_allocator<std::_Sp_counted_ptr_inplace<std::__future_base::_State_baseV2, std::allocator<void>, (__gnu_cxx::_Lock_policy)2> >::allocate(unsigned long, void const*) (new_allocator.h:137) ==45452== by 0x49259A0: allocate (allocator.h:183) ==45452== by 0x49259A0: std::allocator_traits<std::allocator<std::_Sp_counted_ptr_inplace<std::__future_base::_State_baseV2, std::allocator<void>, (__gnu_cxx::_Lock_policy)2> > >::allocate(std::allocator<std::_Sp_counted_ptr_inplace<std::__future_base::_State_baseV2, std::allocator<void>, (__gnu_cxx::_Lock_policy)2> >&, unsigned long) (alloc_traits.h:464) ==45452== by 0x4925840: std::__allocated_ptr<std::allocator<std::_Sp_counted_ptr_inplace<std::__future_base::_State_baseV2, std::allocator<void>, (__gnu_cxx::_Lock_policy)2> > > std::__allocate_guarded<std::allocator<std::_Sp_counted_ptr_inplace<std::__future_base::_State_baseV2, std::allocator<void>, (__gnu_cxx::_Lock_policy)2> > >(std::allocator<std::_Sp_counted_ptr_inplace<std::__future_base::_State_baseV2, std::allocator<void>, (__gnu_cxx::_Lock_policy)2> >&) (allocated_ptr.h:98) ==45452== by 0x4925750: std::__shared_count<(__gnu_cxx::_Lock_policy)2>::__shared_count<std::__future_base::_State_baseV2, std::allocator<void>>(std::__future_base::_State_baseV2*&, std::_Sp_alloc_shared_tag<std::allocator<void> >) (shared_ptr_base.h:969) ==45452== by 0x49256F6: std::__shared_ptr<std::__future_base::_State_baseV2, (__gnu_cxx::_Lock_policy)2>::__shared_ptr<std::allocator<void>>(std::_Sp_alloc_shared_tag<std::allocator<void> >) (shared_ptr_base.h:1712) ==45452== by 0x49256B4: std::shared_ptr<std::__future_base::_State_baseV2>::shared_ptr<std::allocator<void>>(std::_Sp_alloc_shared_tag<std::allocator<void> >) (shared_ptr.h:464) ==45452== by 0x49255F6: std::shared_ptr<std::__future_base::_State_baseV2> std::make_shared<std::__future_base::_State_baseV2>() (shared_ptr.h:1009) ==45452== by 0x4955D08: std::promise<unsigned long>::promise() (future:1080) ==45452== by 0x49549EB: RunOnIoThread(std::function<unsigned long ()>, unsigned long) (schedlib.cc:307) ... ==45452== Other segment start (thread 1) ==45452== at 0x4854114: pthread_mutex_unlock (in /usr/libexec/valgrind/vgpreload_drd-amd64-linux.so) ==45452== by 0x4914602: __gthread_mutex_unlock(pthread_mutex_t*) (gthr-default.h:779) ==45452== by 0x4916D44: std::mutex::unlock() (std_mutex.h:118) ==45452== by 0x4914807: std::lock_guard<std::mutex>::~lock_guard() (std_mutex.h:235) ==45452== by 0x49546FA: PeekNewJobs(int, short, void*) (schedlib.cc:237) ... // this is basically the job processing on IO thread, taking scheduled tasks and running them
核心代码实现
MyValues RunOnIoThread(std::function<MyValues()> act, MyValues onRejection = MyValues::INVOCATION_ERROR) { if (!act) return MyValues::BAD_ARGUMENT; // 如果当前是IO线程,直接执行 if(ThisIsIoThread()) return act(); std::promise<MyValues> pro; try { ScheduleOnIoThread([&]() { try { auto res = act(); pro.set_value(res); } catch (...) { pro.set_exception(std::current_exception()); } }); } catch (...) { return onRejection; } return pro.get_future().get(); // 若有异常则重新抛出 }
原因分析
- DRD对libstdc++内部实现的识别局限:DRD依赖pthread同步原语检测竞争,而libstdc++的
std::promise/future实现采用了自定义的原子futex机制(从报错中的std::__atomic_futex_unsigned可看出),DRD无法识别这种内部同步逻辑,因此误判为数据竞争。 - 手动加锁的反效果:你手动添加的
std::mutex与libstdc++内部的同步机制重叠,DRD会将内部原子操作与你的锁操作视为无关联的竞争行为,反而加剧了误报。
解决办法
1. 添加DRD抑制规则忽略误报
创建一个抑制文件(例如drd_future_suppress.txt),将libstdc++中future相关的内部函数加入白名单:
{ future_wait_race Memcheck:Race fun:load fun:_M_load fun:std::__atomic_futex_unsigned<*>::_M_load_and_test_until fun:std::__atomic_futex_unsigned<*>::_M_load_and_test fun:_M_load_when_equal fun:std::__future_base::_State_baseV2::wait fun:std::__basic_future<*>::_M_get_result fun:std::future<*>::get } { promise_set_race Memcheck:Race fun:store fun:_M_store fun:std::__atomic_futex_unsigned<*>::_M_store fun:std::__future_base::_State_baseV2::_M_set_result fun:std::promise<*>::set_value fun:std::promise<*>::set_exception }
运行Valgrind时指定该抑制文件:
valgrind --tool=drd --suppress=drd_future_suppress.txt ./your_program
2. 用Helgrind替代DRD验证线程安全性
Helgrind是Valgrind的另一个线程检测工具,对C++标准库的支持更完善,通常不会误报std::promise/future的合法操作。运行命令:
valgrind --tool=helgrind ./your_program
3. 确认代码本身的安全性
你的代码逻辑是安全的:
- 栈上的
std::promise对象pro会一直存活到get()返回,而get()会等待IO线程的lambda执行完成,因此lambda访问pro时不会出现悬空引用。 - 若
ScheduleOnIoThread能保证lambda被提交到IO线程执行,那么整个流程的线程同步是正确的。
内容的提问来源于stack exchange,提问作者PasterOfMuppets
相关产品推荐
相关产品推荐

