MIT 6.S081 Barrier多线程实现偶发阻塞问题求助
问题排查:MIT 6.S081 Barrier实验实现问题
我正在完成MIT 6.S081课程的Barrier实验,需要实现barrier函数,让所有线程到达后才能继续执行。我的实现有时成功有时失败,因此在barrier函数中加入了调试打印:printf("%d :%d\n", bstate.nthread, syscall(__NR_gettid));。
成功时输出
0 :81906 0 :81907 0 :81907 0 :81906 0 :81906 1 :81907 0 :81906 1 :81907 0 :81906 0 :81907 0 :81907 0 :81906 0 :81907 0 :81906 0 :81906 1 :81907 0 :81907 0 :81906 0 :81906 1 :81907 OK; passed
失败时输出
失败时程序会卡在最后一行输出后:
0 :81909 0 :81910 0 :81910 1 :81909 0 :81910 0 :81909 0 :81909 0 :81910 0 :81909 1 :81910 0 :81910 1 :81909 0 :81909 1 :81910 0 :81910 1 :81909 0 :81909 1 :81910 0 :81910 1 :81909
实现代码
#include <stdlib.h> #include <unistd.h> #include <stdio.h> #include <assert.h> #include <pthread.h> #include <sys/types.h> #include <unistd.h> #include <sys/syscall.h> static int nthread = 1; static int round = 0; struct barrier { pthread_mutex_t barrier_mutex; pthread_cond_t barrier_cond; int nthread; // Number of threads that have reached this round of the barrier int round; // Barrier round } bstate; static void barrier_init(void) { assert(pthread_mutex_init(&bstate.barrier_mutex, NULL) == 0); assert(pthread_cond_init(&bstate.barrier_cond, NULL) == 0); bstate.nthread = 0; } static void barrier() { // YOUR CODE HERE // // Block until all threads have called barrier() and // then increment bstate.round. // // printf("%d :%d\n", bstate.nthread, syscall(__NR_gettid)); if (++bstate.nthread % nthread == 0) { pthread_cond_broadcast(&bstate.barrier_cond); // wake up every thread sleeping on cond bstate.round++; } else pthread_cond_wait(&bstate.barrier_cond, &bstate.barrier_mutex); // go to sleep on cond, releasing lock mutex, acquiring upon wake up } static void * thread(void *xa) { long n = (long)xa; long delay; int i; for (i = 0; i < 10; i++) { int t = bstate.round; assert(i == t); barrier(); usleep(random() % 100); } // printf("a"); return 0; } int main(int argc, char *argv[]) { pthread_t *tha; void *value; long i; double t1, t0; if (argc < 2) { fprintf(stderr, "%s: %s nthread\n", argv[0], argv[0]); exit(-1); } nthread = atoi(argv[1]); tha = malloc(sizeof(pthread_t) * nthread); srandom(0); barrier_init(); for (i = 0; i < nthread; i++) { assert(pthread_create(&tha[i], NULL, thread, (void *)i) == 0); printf("%d\n", syscall(__NR_gettid)); } for (i = 0; i < nthread; i++) { assert(pthread_join(tha[i], &value) == 0); } printf("OK; passed\n"); }
问题分析与修复
核心问题
未用互斥锁保护共享变量:
bstate.nthread和bstate.round是多线程共享变量,代码读写这些变量时未加锁,导致竞态条件——多个线程同时修改bstate.nthread会出现计数错误,比如两个线程同时读到nthread=0后都自增到1,最终计数永远达不到总线程数。pthread_cond_wait使用错误:调用pthread_cond_wait必须先持有对应的互斥锁,代码直接调用该函数违反了pthread条件变量的使用规则,会触发未定义行为,这是程序卡死的主要原因之一。唤醒后未重置计数且未检查轮次:最后一个线程唤醒其他线程后,未重置
bstate.nthread为0,导致下一轮barrier的计数从之前的数值累加;同时线程被唤醒后未检查当前轮次是否更新,可能出现虚假唤醒(比如其他轮次的唤醒信号被当前轮次的线程接收)。
修复后的barrier函数
static void barrier() { // 加锁保护所有共享变量的访问 pthread_mutex_lock(&bstate.barrier_mutex); // 记录当前轮次,避免虚假唤醒 int current_round = bstate.round; // 线程到达,计数+1 bstate.nthread++; printf("%d :%d\n", bstate.nthread - 1, syscall(__NR_gettid)); // 如果是最后一个到达的线程 if (bstate.nthread == nthread) { // 重置计数,为下一轮做准备 bstate.nthread = 0; // 轮次+1 bstate.round++; // 唤醒所有等待的线程 pthread_cond_broadcast(&bstate.barrier_cond); } else { // 循环等待直到当前轮次结束,避免虚假唤醒 while (bstate.round == current_round) { pthread_cond_wait(&bstate.barrier_cond, &bstate.barrier_mutex); } } // 释放锁 pthread_mutex_unlock(&bstate.barrier_mutex); }
修复说明
- 互斥锁保护:所有对
bstate成员的读写操作都在互斥锁保护下进行,彻底避免竞态条件。 - 正确使用
pthread_cond_wait:调用前先加锁,函数执行时自动释放锁,被唤醒后重新获取锁,符合pthread规范。 - 避免虚假唤醒:用
while循环检查轮次是否更新,确保线程只有在当前轮次真正结束时才继续执行。 - 重置计数:最后一个线程唤醒其他线程前,将
bstate.nthread重置为0,保证下一轮barrier的计数从0开始。
其他优化
- 全局变量
round未被使用,可直接删除,避免混淆。
内容的提问来源于stack exchange,提问作者o o
相关产品推荐
相关产品推荐

