You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

分离线程等待条件变量时的死锁问题:多线程模拟器故障排查

问题背景

我有一个对时间要求极高的模拟器,需要尽可能高效运行。为提升数据吞吐量,我将每个数据帧分配给多个工作线程,为避免每次创建线程的开销,工作线程设为分离状态,永不退出,等待条件变量触发以处理就绪数据。注意:所有线程均操作各自固定且不重叠的内存地址,无数据共享。

每个工作线程被分配独立的thread_data_t结构体,执行以下函数:

void worker(thread_data_t* thread_data) {
    
    thread_data->status = THREAD_STATUS_WAITING;

    while (thread_data->status != THREAD_STATUS_CLOSEING) {

        std::unique_lock<std::mutex> lock(thread_data->mutex);

        // thread_data->status set by calling thread, along with .notify_one()
        thread_data->condition.wait(lock, [thread_data]() { return thread_data->status != THREAD_STATUS_WAITING; });
        
        if (thread_data->status == THREAD_STATUS_ASSIGNED) {
            // 执行计算并设置thread_data->result
        }

        thread_data->status = THREAD_STATUS_WAITING;
    }
}

模拟器函数在线程数据就绪时,通过设置thread_data->status = THREAD_STATUS_ASSIGNED触发线程工作,代码如下:

size_t simulate(simulator_t self) {

    size_t sum = 0;

    while (sum < self->done_count) {
    
        // 通知所有线程数据已准备好,开始工作
        for (size_t i = 0; i < self->num_threads; i++) {

            self->calculation_threads[i].result = 0;

            {
                std::lock_guard<std::mutex> guard(self->calculation_threads[i].mutex);
                self->calculation_threads[i].status = THREAD_STATUS_ASSIGNED;
            }
    
            self->calculation_threads[i].condition.notify_one();
        }

        // 所有线程应已开始工作

        // 等待所有线程完成
        for (size_t i = 0; i < self->num_threads; i++) {

            while (self->calculation_threads[i].status != THREAD_STATUS_WAITING) { 
                // 这里出现了死循环
                self->calculation_threads[i].condition.notify_one(); // 再次通知,防止第一次通知丢失

            }

            sum += self->calculation_threads[i].result;
        }

    }
}

出现的问题

  1. 首次调用simulate()运行正常,但第二次调用时,模拟器卡在等待线程结束的循环中,线程未开始工作。
  2. 线程数≤2时正常,≥3时出现故障。所有thread_data仅在worker或simulate函数中修改。

我怀疑是mutex使用有误,请问问题原因是什么?


问题原因与修复方案

核心问题:状态变量读写未同步,引发竞态条件

你遇到的死循环本质是无锁读取状态变量导致的竞态条件,和mutex的使用直接相关:

  1. 状态读取未加锁,无法保证可见性
    在simulate的等待循环中,你直接无锁读取thread_data->status:

    while (self->calculation_threads[i].status != THREAD_STATUS_WAITING) { ... }
    

    status是多线程共享的变量(worker线程会修改它),无锁读取无法保证simulate线程看到最新值——CPU缓存可能让它一直读到旧的THREAD_STATUS_ASSIGNED状态,直接陷入死循环。

  2. 线程数≥3时问题暴露是调度放大效应
    线程数少的时候,调度逻辑简单,缓存一致性的问题不容易触发,属于偶然的正常表现;线程数增多后,调度不确定性提升,状态不同步的概率急剧升高,直接暴露故障。

  3. 重复通知无意义,反而可能加剧混乱
    循环中重复调用notify_one()无法解决根本问题,因为worker线程可能已经处于WAITING状态,重复通知不会改变它的状态,而simulate线程因为看不到最新状态,会一直死循环。

修复方案:所有状态读写必须加锁

所有对thread_data->status的读写操作,必须在持有对应mutex的前提下进行,同时用条件变量替代忙轮询提升效率:

修改simulate中的等待逻辑:

// 等待所有线程完成
for (size_t i = 0; i < self->num_threads; i++) {
    std::unique_lock<std::mutex> lock(self->calculation_threads[i].mutex);
    // 用条件变量等待状态变为WAITING,避免忙轮询浪费CPU
    self->calculation_threads[i].condition.wait(lock, [&]() {
        return self->calculation_threads[i].status == THREAD_STATUS_WAITING;
    });
    sum += self->calculation_threads[i].result;
}

worker线程的逻辑无需修改,它已经正确在持有锁的情况下修改status,符合线程同步要求。


内容的提问来源于stack exchange,提问作者itzFlubby

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 20:07:18