std::shared_ptr销毁线程安全疑问:为何触发Helgrind数据竞争?
我认为以下C++程序完全安全,因为两个线程操作的是不同的std::shared_ptr实例,但Valgrind的Helgrind工具检测到了数据竞争错误,请问我的理解是否正确?
给定C++程序
#include <iostream> #include <memory> #include <thread> void threadFunction(std::shared_ptr<int> ptr) { std::cout << "Worker thread: " << *ptr << std::endl; } int main() { auto sharedInt = std::make_shared<int>(42); std::cout << "Main thread: " << *sharedInt << std::endl; std::thread t(threadFunction, sharedInt); sharedInt.reset(); t.join(); return 0; }
Helgrind检测到的数据竞争错误
==12797== Possible data race during write of size 4 at 0x4E38088 by thread #1 ==12797== Locks held: none ==12797== at 0x10A4B9: std::_Sp_counted_base<(__gnu_cxx::_Lock_policy)2>::_M_release() (shared_ptr_base.h:343) ==12797== by 0x10A794: std::__shared_count<(__gnu_cxx::_Lock_policy)2>::~__shared_count() (shared_ptr_base.h:1071) ==12797== by 0x10A663: std::__shared_ptr<int, (__gnu_cxx::_Lock_policy)2>::~__shared_ptr() (shared_ptr_base.h:1524) ==12797== by 0x10A8EF: std::__shared_ptr<int, (__gnu_cxx::_Lock_policy)2>::reset() (shared_ptr_base.h:1642) ==12797== by 0x10A300: main (shared_ptr.cpp:17) ==12797== ==12797== This conflicts with a previous read of size 4 by thread #2 ==12797== Locks held: none ==12797== at 0x10A49E: std::_Sp_counted_base<(__gnu_cxx::_Lock_policy)2>::_M_release() (shared_ptr_base.h:337) ==12797== by 0x10A794: std::__shared_count<(__gnu_cxx::_Lock_policy)2>::~__shared_count() (shared_ptr_base.h:1071) ==12797== by 0x10A663: std::__shared_ptr<int, (__gnu_cxx::_Lock_policy)2>::~__shared_ptr() (shared_ptr_base.h:1524) ==12797== by 0x10A67F: std::shared_ptr<int>::~shared_ptr() (shared_ptr.h:175) ==12797== by 0x10BB2D: void std::__invoke_impl<void, void (*)(std::shared_ptr<int>), std::shared_ptr<int> >(std::__invoke_other, void (*&&)(std::shared_ptr<int>), std::shared_ptr<int>&&) (invoke.h:61) ==12797== by 0x10BA74: std::__invoke_result<void (*)(std::shared_ptr<int>), std::shared_ptr<int> >::type std::__invoke<void (*)(std::shared_ptr<int>), std::shared_ptr<int> >(void (*&&)(std::shared_ptr<int>), std::shared_ptr<int>&&) (invoke.h:96) ==12797== by 0x10B9E4: void std::thread::_Invoker<std::tuple<void (*)(std::shared_ptr<int>), std::shared_ptr<int> > >::_M_invoke<0ul, 1ul>(std::_Index_tuple<0ul, 1ul>) (std_thread.h:292) ==12797== by 0x10B985: std::thread::_Invoker<std::tuple<void (*)(std::shared_ptr<int>), std::shared_ptr<int> > >::operator()() (std_thread.h:299) ==12797== Address 0x4e38088 is 8 bytes inside a block of size 24 alloc'd ==12797== at 0x4846023: operator new(unsigned long) (vg_replace_malloc.c:483) ==12797== by 0x10B51E: std::__new_allocator<std::_Sp_counted_ptr_inplace<int, std::allocator<void>, (__gnu_cxx::_Lock_policy)2> >::allocate(unsigned long, void const*) (new_allocator.h:147) ==12797== by 0x10B10F: allocate (alloc_traits.h:482) ==12797== by 0x10B10F: std::__allocated_ptr<std::allocator<std::_Sp_counted_ptr_inplace<int, std::allocator<void>, (__gnu_cxx::_Lock_policy)2> > > std::__allocate_guarded<std::allocator<std::_Sp_counted_ptr_inplace<int, std::allocator<void>, (__gnu_cxx::_Lock_policy)2> > >(std::allocator<std::_Sp_counted_ptr_inplace<int, std::allocator<void>, (__gnu_cxx::_Lock_policy)2> >&) (allocated_ptr.h:98) ==12797== by 0x10AF48: std::__shared_count<(__gnu_cxx::_Lock_policy)2>::__shared_count<int, std::allocator<void>, int>(int*&, std::_Sp_alloc_shared_tag<std::allocator<void> >, int&&) (shared_ptr_base.h:969) ==12797== by 0x10AD29: std::__shared_ptr<int, (__gnu_cxx::_Lock_policy)2>::__shared_ptr<std::allocator<void>, int>(std::_Sp_alloc_shared_tag<std::allocator<void> >, int&&) (shared_ptr_base.h:1712) ==12797== by 0x10AA60: std::shared_ptr<int>::shared_ptr<std::allocator<void>, int>(std::_Sp_alloc_shared_tag<std::allocator<void> >, int&&) (shared_ptr.h:464) ==12797== by 0x10A753: std::shared_ptr<std::enable_if<!std::is_array<int>::value, int>::type> std::make_shared<int, int>(int&&) (shared_ptr.h:1010) ==12797== by 0x10A294: main (shared_ptr.cpp:11) ==12797== Block was alloc'd by thread #1
解答
你的理解不正确,问题出在两个std::shared_ptr实例共享的引用计数对象上:
- 创建线程
t并传入sharedInt时,会触发shared_ptr的拷贝,此时引用计数从1变为2。 - 主线程调用
sharedInt.reset()时,会减少引用计数(从2变为1),这个操作会修改引用计数对象的内存。 - 工作线程的
ptr参数在函数结束时销毁,同样会减少引用计数(从1变为0,最终释放托管对象),这个操作也会访问引用计数对象的内存。
这两个操作(主线程的写、工作线程的读/写)是并发执行的,且没有同步机制,因此触发了数据竞争。
需要注意的是,C++标准规定shared_ptr的引用计数更新是原子操作,但Helgrind报错的原因可能有两种:
- 你的编译器实现中,
_M_release()操作包含了非原子的读步骤(比如先读取当前计数,再判断是否需要释放),这个读操作和其他线程的写操作发生了竞争。 - Helgrind对某些原子操作的识别存在偏差,或者你的实现没有用完全原子的方式处理引用计数的所有环节。
要修复这个问题,只需确保主线程在调用reset()前等待工作线程完成,也就是把sharedInt.reset()移到t.join()之后:
int main() { auto sharedInt = std::make_shared<int>(42); std::cout << "Main thread: " << *sharedInt << std::endl; std::thread t(threadFunction, sharedInt); t.join(); // 先等待线程执行完毕 sharedInt.reset(); return 0; }
这样就能避免引用计数的并发访问,消除数据竞争。
内容的提问来源于stack exchange,提问作者Michał Walenciak
相关产品推荐
相关产品推荐

