为何基于pthreads的多线程程序比单线程程序运行更慢?
问题:多线程版本为何比单线程版本运行速度慢?
我是多线程编程新手,编写了一个计算0到10000的平方值并存入数组的程序。但单线程版本的运行速度远快于我用8个线程(我的机器配备8核)实现的并行版本。附上两段代码,想请教出现这种情况的原因:
单线程程序代码
/*以下是单线程程序:*/ #define ARRAYSIZE 10000 int main(void) { int array[ARRAYSIZE]; int i; for (i=0; i<ARRAYSIZE; i++) { array[i]=i*i; } return 0; }
并行计算程序代码
/*以下是并行计算程序*/ #include <stdio.h> #include <pthread.h> #define ARRAYSIZE 10000 #define NUMTHREADS 8 /*因我的机器有8核*/ struct ThreadData { int start; int stop; int* array; }; void* squarer (struct ThreadData* td); /* 将i²存入数组中索引从start到stop-1的位置 */ void* squarer (struct ThreadData* td) { struct ThreadData* data = (struct ThreadData*) td; int start=data->start; int stop=data->stop; int* array=data->array; int i; for(i= start; i<stop; i++) { array[i]=i*i; } return NULL; } int main(void) { int array[ARRAYSIZE]; pthread_t thread[NUMTHREADS]; struct ThreadData data[NUMTHREADS]; int i; int tasksPerThread= (ARRAYSIZE + NUMTHREADS - 1)/ NUMTHREADS; /* 为线程分配任务,准备参数 */ /* 即在此示例中,我将循环划分为8个区域:0..1250、1250..2500等,2500..3750 */ for(i=0; i<NUMTHREADS;i++) { data[i].start=i*tasksPerThread; data[i].stop=(i+1)*tasksPerThread; data[i].array=array; data[NUMTHREADS-1].stop=ARRAYSIZE; } for(i=0; i<NUMTHREADS;i++) { pthread_create(&thread[i], NULL, squarer, &data[i]); } for(i=0; i<NUMTHREADS;i++) { pthread_join(thread[i], NULL); } return 0; }
原因分析
线程管理开销远超计算收益:你的计算任务过于简单——每个线程仅需完成1250次乘法和数组赋值操作。而创建线程、线程调度、调用
pthread_join等待线程结束这些操作的开销,远大于并行计算能节省的时间。单线程无需这些额外开销,自然运行更快。缓存一致性同步成本:所有线程都在写入同一个数组,尽管每个线程操作不同区域,但CPU缓存一致性协议(如MESI)需要同步各核心的缓存状态,这会产生额外开销。单线程操作数组时,缓存可持续命中,没有这类同步成本。
任务划分粒度太细:10000个元素分给8个线程,每个线程仅处理1250个元素,这种细粒度划分让并行的优势完全被线程管理开销抵消。若将数组规模放大到百万甚至千万级别,多线程版本的性能优势会显现出来。
编译器优化差异:单线程代码结构简单,编译器可进行更激进的优化(如循环展开、寄存器重排);而多线程代码涉及线程间内存访问,编译器的优化空间受限,进一步拉大了性能差距。
内容的提问来源于stack exchange,提问作者JohnDoe
相关产品推荐
相关产品推荐

