You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Folly的rcu_domain.retire需用非阻塞half_sync?求死锁示例

Folly Rcu.h中half_sync(true)的死锁场景示例

背景与问题

我正在研究Folly的Rcu.h实现,重点关注retire函数里half_sync的使用逻辑,相关核心代码如下:

void retire(list_node* node) noexcept {
 // ...
 if (...) {
   list_head finished;
   {
     std::lock_guard<std::mutex> g(syncMutex_);
     half_sync(false, finished);
   }
   // ...
 }
}

void half_sync(bool blocking, list_head& finished) {
  auto curr = version_.load(std::memory_order_acquire);
  auto next = curr + 1;
    
  // ...
  if (blocking) {
    counters_.waitForZero(next & 1);
  } else {
    if (!counters_.epochIsClear(next & 1)) {
      return;
    }
  }
  // ...
}

我对half_sync的理解是:它会根据blocking参数,要么等待、要么检查所有读端临界区进入当前epoch,之后递增version_,并把两个epoch前待回收的项加入finished列表,供后续处理。

retire函数里调用的是非阻塞版本half_sync(false),代码注释里明确警告:

Note that it's likely we hold a read lock here, so we can only half_sync(false). half_sync(true) or a synchronize() call might block forever.

我测试时用half_sync(true)没碰到问题,想知道具体什么场景会触发死锁?

死锁场景示例

下面是一个能复现死锁的极简示例:

#include <folly/synchronization/Rcu.h>
#include <thread>
#include <mutex>

folly::rcu_domain rcu;
std::mutex resource_mutex;
int* shared_data = new int(42);

// 读端线程:持有RCU读锁+资源锁
void reader_thread() {
  folly::rcu_read_lock guard(rcu);
  std::lock_guard<std::mutex> lock(resource_mutex);
  
  // 访问共享数据
  printf("Shared data: %d\n", *shared_data);
  
  // 故意延长持有时间,让写端触发retire并调用half_sync(true)
  std::this_thread::sleep_for(std::chrono::seconds(2));
}

// 写端线程:修改共享数据并触发retire
void writer_thread() {
  // 先拿资源锁,确保读端的RCU读锁先被持有
  std::lock_guard<std::mutex> lock(resource_mutex);
  
  int* old_data = shared_data;
  shared_data = new int(100);
  
  // 调用retire,此时如果把内部的half_sync(false)改成half_sync(true)
  rcu.retire(old_data, [](void* p) { delete static_cast<int*>(p); });
}

int main() {
  std::thread t1(reader_thread);
  // 等待读端先进入临界区,持有RCU读锁和资源锁
  std::this_thread::sleep_for(std::chrono::seconds(1));
  std::thread t2(writer_thread);
  
  t1.join();
  t2.join();
  delete shared_data;
  return 0;
}

死锁原因解释

  1. 读端线程:先获取RCU读锁,再拿到resource_mutex,然后进入睡眠,同时持有这两个锁。
  2. 写端线程:先拿到resource_mutex,然后尝试调用retire——如果retire内部调用half_sync(true),会触发counters_.waitForZero(next & 1),等待所有读端退出当前RCU epoch。
  3. 但读端线程因为持有resource_mutex,被写端线程的锁持有顺序卡住,无法退出RCU读锁;而写端线程又卡在half_sync(true)等待读端退出,形成循环等待:
    • 读端:持有RCU读锁 → 等待写端释放resource_mutex(但写端不会放)
    • 写端:持有resource_mutex → 等待读端释放RCU读锁(但读端不会放)

这就是注释里说的死锁场景——当retire被调用时,当前线程可能间接持有了某个锁,导致读端线程无法退出RCU临界区,最终half_sync(true)会无限阻塞。

内容的提问来源于stack exchange,提问作者RSIMB GO

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 20:52:10