使用原子变量的事件驱动程序意外卡顿问题排查求助
问题描述
基于线程池与原子变量实现的事件驱动架构,通过原子变量ntasks实现栅栏事件作为同步点:begin_event()递增ntasks,end_event()递减,当计数为0时通知事件循环继续。但程序在特定条件下会卡在loop_cv.wait()处;使用调试器或添加输出语句时问题无法重现,将等待逻辑替换为自旋锁后卡顿消失。运行环境为Windows系统、MSVC编译器、Intel i7-10750H CPU,线程池采用第三方实现。
简化版代码如下:
#include <cstdint> #include <atomic> #include <mutex> #include <thread> #include <queue> #include <cstdio> #include "ThreadPool.h" enum event_t { update, render, fence = 0xffffffff }; std::mutex queue_mutex; std::queue<event_t> queue; std::atomic<long long> ntasks; ThreadPool thread_pool(4); std::mutex loop_mutex; std::condition_variable loop_cv; void publish_event(event_t t) { std::unique_lock<std::mutex> lock(queue_mutex); queue.push(t); } bool try_get_event(event_t& x) { std::unique_lock<std::mutex> lock(queue_mutex); if (queue.empty()) return false; x = queue.front(); queue.pop(); return true; } void onUpdate() { printf("%s", "Update!"); publish_event(render); publish_event(fence); // Wait until render event finished publish_event(update); // Start the next pass } void onRender() { // Pass } // Called before event execution start void begin_event() { ntasks.fetch_add(1, std::memory_order_release); // Increase the count of excution tasks } void end_event() { if (ntasks.fetch_sub(1, std::memory_order_acq_rel) == 1) // Decrease the count of excution tasks loop_cv.notify_one(); // If there are no more tasks, we can notify to release the fence } int main() { publish_event(update); // Initial event to start the event loop // Event Loop while (true) { event_t event; if (try_get_event(event)) { switch (event) { case update: begin_event(); thread_pool.enqueue([] { onUpdate(); end_event(); }); break; case render: begin_event(); thread_pool.enqueue([] { onRender(); end_event(); }); break; case fence: { std::unique_lock<std::mutex> lock(loop_mutex); // Wait until there are no more tasks loop_cv.wait(lock, [] { return ntasks.load(std::memory_order_acquire) <= 0; }); } break; } } } }
卡顿原因分析
- 条件变量唤醒丢失:这是核心问题。
end_event()中调用loop_cv.notify_one()时未持有loop_mutex,存在致命时序窗口:主线程在loop_cv.wait()的谓词检查阶段发现ntasks>0,刚释放loop_mutex准备进入休眠,此时线程池执行end_event()将ntasks减至0并调用notify_one(),该通知会直接丢失,主线程进入休眠后再也无法被唤醒。调试或添加输出语句时,额外的同步操作改变了执行时序,避开了这个窗口;自旋锁通过轮询ntasks无需依赖通知,因此不会卡顿。 - 事件循环忙等待:主线程通过
while(true)不断调用try_get_event()轮询事件队列,无事件时持续占用CPU,属于性能问题但并非卡顿直接原因。
改进方案
1. 修复条件变量唤醒丢失问题
调用loop_cv.notify_one()前必须持有loop_mutex,确保通知不会丢失,同时保证ntasks的修改对主线程可见:
void end_event() { if (ntasks.fetch_sub(1, std::memory_order_acq_rel) == 1) { std::lock_guard<std::mutex> lock(loop_mutex); loop_cv.notify_one(); } }
2. 优化事件循环的忙等待
为事件队列添加条件变量,让主线程在队列空时休眠,避免无意义的CPU占用:
// 新增队列条件变量 std::condition_variable queue_cv; void publish_event(event_t t) { std::unique_lock<std::mutex> lock(queue_mutex); queue.push(t); queue_cv.notify_one(); // 通知主线程有新事件 } // 修改事件循环 int main() { publish_event(update); while (true) { event_t event; std::unique_lock<std::mutex> queue_lock(queue_mutex); // 等待队列非空 queue_cv.wait(queue_lock, []{ return !queue.empty(); }); event = queue.front(); queue.pop(); queue_lock.unlock(); switch (event) { case update: begin_event(); thread_pool.enqueue([] { onUpdate(); end_event(); }); break; case render: begin_event(); thread_pool.enqueue([] { onRender(); end_event(); }); break; case fence: { std::unique_lock<std::mutex> loop_lock(loop_mutex); loop_cv.wait(loop_lock, [] { return ntasks.load(std::memory_order_acquire) <= 0; }); } break; } } }
3. 内存顺序优化(可选)
begin_event()中fetch_add使用std::memory_order_relaxed即可,因为只需要保证计数的原子性,无需同步其他内存操作;end_event()的fetch_sub保持std::memory_order_acq_rel确保操作的可见性:
void begin_event() { ntasks.fetch_add(1, std::memory_order_relaxed); }
内容的提问来源于stack exchange,提问作者X1ngChui
相关产品推荐
相关产品推荐

