分离线程等待条件变量时的死锁问题:多线程模拟器故障排查
我有一个对时间要求极高的模拟器,需要尽可能高效运行。为提升数据吞吐量,我将每个数据帧分配给多个工作线程,为避免每次创建线程的开销,工作线程设为分离状态,永不退出,等待条件变量触发以处理就绪数据。注意:所有线程均操作各自固定且不重叠的内存地址,无数据共享。
每个工作线程被分配独立的thread_data_t结构体,执行以下函数:
void worker(thread_data_t* thread_data) { thread_data->status = THREAD_STATUS_WAITING; while (thread_data->status != THREAD_STATUS_CLOSEING) { std::unique_lock<std::mutex> lock(thread_data->mutex); // thread_data->status set by calling thread, along with .notify_one() thread_data->condition.wait(lock, [thread_data]() { return thread_data->status != THREAD_STATUS_WAITING; }); if (thread_data->status == THREAD_STATUS_ASSIGNED) { // 执行计算并设置thread_data->result } thread_data->status = THREAD_STATUS_WAITING; } }
模拟器函数在线程数据就绪时,通过设置thread_data->status = THREAD_STATUS_ASSIGNED触发线程工作,代码如下:
size_t simulate(simulator_t self) { size_t sum = 0; while (sum < self->done_count) { // 通知所有线程数据已准备好,开始工作 for (size_t i = 0; i < self->num_threads; i++) { self->calculation_threads[i].result = 0; { std::lock_guard<std::mutex> guard(self->calculation_threads[i].mutex); self->calculation_threads[i].status = THREAD_STATUS_ASSIGNED; } self->calculation_threads[i].condition.notify_one(); } // 所有线程应已开始工作 // 等待所有线程完成 for (size_t i = 0; i < self->num_threads; i++) { while (self->calculation_threads[i].status != THREAD_STATUS_WAITING) { // 这里出现了死循环 self->calculation_threads[i].condition.notify_one(); // 再次通知,防止第一次通知丢失 } sum += self->calculation_threads[i].result; } } }
出现的问题
- 首次调用
simulate()运行正常,但第二次调用时,模拟器卡在等待线程结束的循环中,线程未开始工作。 - 线程数≤2时正常,≥3时出现故障。所有
thread_data仅在worker或simulate函数中修改。
我怀疑是mutex使用有误,请问问题原因是什么?
核心问题:状态变量读写未同步,引发竞态条件
你遇到的死循环本质是无锁读取状态变量导致的竞态条件,和mutex的使用直接相关:
状态读取未加锁,无法保证可见性
在simulate的等待循环中,你直接无锁读取thread_data->status:while (self->calculation_threads[i].status != THREAD_STATUS_WAITING) { ... }status是多线程共享的变量(worker线程会修改它),无锁读取无法保证simulate线程看到最新值——CPU缓存可能让它一直读到旧的THREAD_STATUS_ASSIGNED状态,直接陷入死循环。线程数≥3时问题暴露是调度放大效应
线程数少的时候,调度逻辑简单,缓存一致性的问题不容易触发,属于偶然的正常表现;线程数增多后,调度不确定性提升,状态不同步的概率急剧升高,直接暴露故障。重复通知无意义,反而可能加剧混乱
循环中重复调用notify_one()无法解决根本问题,因为worker线程可能已经处于WAITING状态,重复通知不会改变它的状态,而simulate线程因为看不到最新状态,会一直死循环。
修复方案:所有状态读写必须加锁
所有对thread_data->status的读写操作,必须在持有对应mutex的前提下进行,同时用条件变量替代忙轮询提升效率:
修改simulate中的等待逻辑:
// 等待所有线程完成 for (size_t i = 0; i < self->num_threads; i++) { std::unique_lock<std::mutex> lock(self->calculation_threads[i].mutex); // 用条件变量等待状态变为WAITING,避免忙轮询浪费CPU self->calculation_threads[i].condition.wait(lock, [&]() { return self->calculation_threads[i].status == THREAD_STATUS_WAITING; }); sum += self->calculation_threads[i].result; }
worker线程的逻辑无需修改,它已经正确在持有锁的情况下修改status,符合线程同步要求。
内容的提问来源于stack exchange,提问作者itzFlubby

