为何Folly的rcu_domain.retire需用非阻塞half_sync?求死锁示例
Folly Rcu.h中
half_sync(true)的死锁场景示例 背景与问题
我正在研究Folly的Rcu.h实现,重点关注retire函数里half_sync的使用逻辑,相关核心代码如下:
void retire(list_node* node) noexcept { // ... if (...) { list_head finished; { std::lock_guard<std::mutex> g(syncMutex_); half_sync(false, finished); } // ... } } void half_sync(bool blocking, list_head& finished) { auto curr = version_.load(std::memory_order_acquire); auto next = curr + 1; // ... if (blocking) { counters_.waitForZero(next & 1); } else { if (!counters_.epochIsClear(next & 1)) { return; } } // ... }
我对half_sync的理解是:它会根据blocking参数,要么等待、要么检查所有读端临界区进入当前epoch,之后递增version_,并把两个epoch前待回收的项加入finished列表,供后续处理。
retire函数里调用的是非阻塞版本half_sync(false),代码注释里明确警告:
Note that it's likely we hold a read lock here, so we can only
half_sync(false).half_sync(true)or asynchronize()call might block forever.
我测试时用half_sync(true)没碰到问题,想知道具体什么场景会触发死锁?
死锁场景示例
下面是一个能复现死锁的极简示例:
#include <folly/synchronization/Rcu.h> #include <thread> #include <mutex> folly::rcu_domain rcu; std::mutex resource_mutex; int* shared_data = new int(42); // 读端线程:持有RCU读锁+资源锁 void reader_thread() { folly::rcu_read_lock guard(rcu); std::lock_guard<std::mutex> lock(resource_mutex); // 访问共享数据 printf("Shared data: %d\n", *shared_data); // 故意延长持有时间,让写端触发retire并调用half_sync(true) std::this_thread::sleep_for(std::chrono::seconds(2)); } // 写端线程:修改共享数据并触发retire void writer_thread() { // 先拿资源锁,确保读端的RCU读锁先被持有 std::lock_guard<std::mutex> lock(resource_mutex); int* old_data = shared_data; shared_data = new int(100); // 调用retire,此时如果把内部的half_sync(false)改成half_sync(true) rcu.retire(old_data, [](void* p) { delete static_cast<int*>(p); }); } int main() { std::thread t1(reader_thread); // 等待读端先进入临界区,持有RCU读锁和资源锁 std::this_thread::sleep_for(std::chrono::seconds(1)); std::thread t2(writer_thread); t1.join(); t2.join(); delete shared_data; return 0; }
死锁原因解释
- 读端线程:先获取RCU读锁,再拿到
resource_mutex,然后进入睡眠,同时持有这两个锁。 - 写端线程:先拿到
resource_mutex,然后尝试调用retire——如果retire内部调用half_sync(true),会触发counters_.waitForZero(next & 1),等待所有读端退出当前RCU epoch。 - 但读端线程因为持有
resource_mutex,被写端线程的锁持有顺序卡住,无法退出RCU读锁;而写端线程又卡在half_sync(true)等待读端退出,形成循环等待:- 读端:持有RCU读锁 → 等待写端释放
resource_mutex(但写端不会放) - 写端:持有
resource_mutex→ 等待读端释放RCU读锁(但读端不会放)
- 读端:持有RCU读锁 → 等待写端释放
这就是注释里说的死锁场景——当retire被调用时,当前线程可能间接持有了某个锁,导致读端线程无法退出RCU临界区,最终half_sync(true)会无限阻塞。
内容的提问来源于stack exchange,提问作者RSIMB GO
相关产品推荐
相关产品推荐

