基于atomic_bool与atomic_int的线程屏障等待实现死锁问题排查
问题描述
我实现了一个等待所有线程到达后执行特定操作的方法,采用atomic_bool与atomic_int实现屏障与等待逻辑,但运行时出现死锁:所有16个线程均阻塞在waitForAllThread函数中,此时atomic_flag_guard为true,但atomic_counter的值仅为2(预期应为15)。编译命令为/pkgs/gccv9.3.0p4/bin/g++ thread_barrier_impl.cpp -std=gnu++17 -lpthread,多次测试均会出现挂起现象,相关代码如下:
#include <iostream> #include <pthread.h> #include <stdio.h> #include <atomic> #include <thread> #define TOTAL_THREADS 16 using namespace std; volatile atomic_bool atomic_flag_guard = 0; // 0号线程根据处理进度决定是否锁定 volatile atomic_int atomic_counter = 0; // 记录已到达的线程数 volatile atomic_int all_processed_object_count = 0; // 已处理对象数 volatile atomic_int total_number_thread = TOTAL_THREADS ; struct thread_local_data { int t_index = 0; int total_thread = 0; }; void waitForAllThread(int thread_index, int &stop_index, size_t interval) { if(thread_index == 0 && all_processed_object_count >= stop_index) { atomic_flag_guard.store(true); printf("guard up\n"); while(atomic_counter != total_number_thread - 1); // 执行任务 atomic_counter.store(0); stop_index += interval; atomic_flag_guard.store(false); } else if( thread_index != 0 && atomic_flag_guard.load()) { ++atomic_counter; printf("arrived %d total arrived %d\n", thread_index, atomic_counter.load()); while(atomic_flag_guard.load()); } } void* thread_processing(void *thread_data) { thread_local_data *tld = static_cast<thread_local_data*>(thread_data); int stop_index = 100; int interval = stop_index; for(int i = 0 ; i < 1000; ++i) { std::this_thread::sleep_for(2ms); ++all_processed_object_count; waitForAllThread(tld->t_index, stop_index, interval); } printf("Exit thread %d\n", tld->t_index); --total_number_thread; return nullptr; } int main() { thread_local_data tld[TOTAL_THREADS ]; for(int i =0; i < TOTAL_THREADS ; ++i){tld[i].total_thread = TOTAL_THREADS ; tld[i].t_index = i;} pthread_t threads[TOTAL_THREADS ]; for (int i = 0; i < TOTAL_THREADS ; i++) { auto *obj = &tld[i]; void *userData = static_cast<void*>(obj); pthread_create(threads + i, NULL, thread_processing, userData); } for (int i = 0; i < TOTAL_THREADS ; i++) { pthread_join(threads[i], NULL); } return 0; }
死锁原因分析
非0号线程的屏障触发条件存在竞态:非0号线程仅在
atomic_flag_guard为true时才会递增atomic_counter,但0号线程设置flag为true的时机,可能晚于部分非0号线程执行waitForAllThread的判断逻辑(此时flag仍为false),这部分线程会直接跳过计数步骤。0号线程会一直等待atomic_counter达到total_number_thread-1(15),但实际只有少数线程完成计数,最终所有线程陷入阻塞。线程局部的stop_index导致同步失效:每个线程的
stop_index是栈上的局部变量,只有0号线程的stop_index会在屏障触发后更新,其他线程的stop_index始终保持初始值100。当all_processed_object_count超过100后,非0号线程每次进入waitForAllThread都会检查flag,但如果0号线程已经将flag重置为false,它们就不会参与计数;而0号线程下次触发flag时,之前错过计数的线程可能再次错过,导致计数永远无法达标。原子变量冗余的volatile修饰:C++标准中
atomic类型本身已经保证内存可见性和原子操作语义,额外添加volatile不仅冗余,还可能干扰编译器对原子操作的优化,甚至引发内存顺序相关的潜在问题。线程退出时修改total_number_thread的时机不合理:线程在完成1000次循环后才递减
total_number_thread,但如果循环过程中存在线程异常退出(逻辑上的可能性),0号线程的等待条件atomic_counter != total_number_thread -1会动态变化,进一步加剧计数逻辑的混乱。
内容的提问来源于stack exchange,提问作者Roman

