《C++ Concurrency in Action》工作窃取线程池线程安全问题咨询
我正在阅读《C++ Concurrency in Action(第二版)》,书中展示了如下基于工作窃取的线程池实现:
// Listing 9.7 Lock-based queue for work stealing class work_stealing_queue { private: typedef function_wrapper data_type; std::deque<data_type> the_queue; mutable std::mutex the_mutex; public: work_stealing_queue() {} work_stealing_queue(const work_stealing_queue& other)=delete; work_stealing_queue& operator=(const work_stealing_queue& other)=delete; void push(data_type data) { std::lock_guard<std::mutex> lock(the_mutex); the_queue.push_front(std::move(data)); } bool empty() const { std::lock_guard<std::mutex> lock(the_mutex); return the_queue.empty(); } bool try_pop(data_type& res) { std::lock_guard<std::mutex> lock(the_mutex); if(the_queue.empty()) { return false; } res=std::move(the_queue.front()); the_queue.pop_front(); return true; } bool try_steal(data_type& res) { std::lock_guard<std::mutex> lock(the_mutex); if(the_queue.empty()) { return false; } res=std::move(the_queue.back()); the_queue.pop_back(); return true; } }; // Listing 9.8 A thread pool that uses work stealing class thread_pool { typedef function_wrapper task_type; std::atomic_bool done; threadsafe_queue<task_type> pool_work_queue; std::vector<std::unique_ptr<work_stealing_queue>> queues; std::vector<std::thread> threads; join_threads joiner; static thread_local work_stealing_queue* local_work_queue; static thread_local unsigned my_index; void worker_thread(unsigned my_index_) { my_index=my_index_; local_work_queue=queues[my_index].get(); while(!done) { run_pending_task(); } } bool pop_task_from_local_queue(task_type& task) { return local_work_queue && local_work_queue->try_pop(task); } bool pop_task_from_pool_queue(task_type& task) { return pool_work_queue.try_pop(task); } bool pop_task_from_other_thread_queue(task_type& task) { for(unsigned i=0;i<queues.size();++i) { unsigned const index=(my_index+i+1)%queues.size(); if(queues[index]->try_steal(task)) { return true; } } return false; } public: thread_pool(): done(false),joiner(threads) { unsigned const thread_count=std::thread::hardware_concurrency(); try { for(unsigned i=0;i<thread_count;++i) { queues.push_back(std::unique_ptr<work_stealing_queue>( new work_stealing_queue)); } for(unsigned i=0;i<thread_count;++i) { threads.push_back(std::thread(&thread_pool::worker_thread,this,i)); } } catch(...) { done=true; throw; } } ~thread_pool() { done=true; } template<typename FunctionType> std::future<typename std::result_of<FunctionType()>::type> submit( FunctionType f) { typedef typename std::result_of<FunctionType()>::type result_type; std::packaged_task<result_type()> task(f); std::future<result_type> res(task.get_future()); if(local_work_queue) { local_work_queue->push(std::move(task)); } else { pool_work_queue.push(std::move(task)); } return res; } void run_pending_task() { task_type task; if(pop_task_from_local_queue(task) || pop_task_from_pool_queue(task) || pop_task_from_other_thread_queue(task)) { task(); } else { std::this_thread::yield(); } } };
问题
在thread_pool的构造函数中,先完成所有work_stealing_queue的构造,再创建工作线程。当工作线程执行到run_pending_task时,会访问thread_pool::queues成员变量。是否可能因内存重排序,导致工作线程访问时queues内元素的构造尚未完成?若不会,其顺序是如何保证的?我未找到这些事件间的synchronize-with关系,恳请解释上述线程安全问题。
解答
不会出现内存重排序导致的未构造完成问题,核心保障来自C++标准对线程创建的内存同步语义:
1. 线程构造的happens-before关系
C++标准明确规定:主线程中在构造std::thread对象之前完成的所有内存操作,都happens-before新线程中执行的任何操作。也就是说,主线程先填充queues容器(完成所有work_stealing_queue实例的构造,并将指针存入容器),再创建工作线程的顺序,会被编译器和CPU严格遵守,不会出现重排序。
2. 隐式的synchronize-with同步点
std::thread的构造函数本身就是一个同步点:主线程构造std::thread前的所有写操作,与新线程的执行起始点之间存在synchronize-with关系。这相当于隐式的内存屏障,确保工作线程在访问queues时,容器内的所有work_stealing_queue实例都已完全构造,不存在未初始化内存访问的风险。
3. 代码层面的具体保证
在thread_pool构造函数中,我们先通过循环完成queues的填充,确保每个work_stealing_queue都已完成初始化,对应的unique_ptr也已正确存入容器。之后才调用std::thread创建工作线程——这一步的同步语义直接保证了工作线程看到的queues是完全初始化后的状态。
另外需要注意,done作为std::atomic_bool,其原子操作提供了额外的内存可见性保证,但这里的核心顺序保障还是来自线程创建的语义,done主要用于控制工作线程的退出逻辑。
内容的提问来源于stack exchange,提问作者pan64271

