短任务并行性能异常:POSIX自旋线程耗时疑问
问题描述
我需要百万次重复并行执行一段微秒级的短代码,因为创建POSIX线程开销太高,所以采用了自旋工作线程的方案。在简化案例里并行运行calc_x和calc_y,原本期望总耗时接近单任务的耗时,但实际结果没达到预期:当nloops=100时并行几乎没收益,nloops=1000时仍比预期慢1/3,而且设置atomic_bool变量trun的耗时还会随nloops增加而上升,想知道这一现象的原因。
代码示例
#include <pthread.h> #include <time.h> #include <stdatomic.h> #include <stdio.h> static int nloops = 100; static int x; static int y; static pthread_t tid; static atomic_bool trun; static atomic_bool texit; static int cnt = 0; static const int MILLION = 1000000; void calc_x() { x = 0; for (int i = 0; i < nloops; ++i) { x += i; } } void calc_y() { y = 0; for (int i = 0; i < nloops; ++i) { y += i; } } void *worker() { while (1) { if (trun) { calc_x(); trun = false; } if (texit) { break; } } return NULL; } float run_x() { struct timespec start, end; clock_gettime(CLOCK_MONOTONIC, &start); for (int i = 0; i < MILLION; ++i) { calc_x(); } clock_gettime(CLOCK_MONOTONIC, &end); return (end.tv_sec - start.tv_sec) + (end.tv_nsec - start.tv_nsec)*0.000000001; } float run_y() { struct timespec start, end; clock_gettime(CLOCK_MONOTONIC, &start); for (int i = 0; i < MILLION; ++i) { calc_y(); } clock_gettime(CLOCK_MONOTONIC, &end); return (end.tv_sec - start.tv_sec) + (end.tv_nsec - start.tv_nsec)*0.000000001; } float run_xy() { struct timespec start, end; clock_gettime(CLOCK_MONOTONIC, &start); for (int i = 0; i < MILLION; ++i) { calc_x(); calc_y(); } clock_gettime(CLOCK_MONOTONIC, &end); return (end.tv_sec - start.tv_sec) + (end.tv_nsec - start.tv_nsec)*0.000000001; } float run_xy_parallel() { struct timespec start, end; clock_gettime(CLOCK_MONOTONIC, &start); for (int i = 0; i < MILLION; ++i) { trun = true; calc_y(); while(trun) ++cnt; } clock_gettime(CLOCK_MONOTONIC, &end); return (end.tv_sec - start.tv_sec) + (end.tv_nsec - start.tv_nsec)*0.000000001; } int main() { trun = false; texit = false; pthread_create(&tid, NULL, &worker, NULL); printf("run_x: %.4f s\n", run_x()); printf("run_y: %.4f s\n", run_y()); printf("run_xy: %.4f s\n", run_xy()); printf("run_xy_parallel: %.4f s\n", run_xy_parallel()); printf("average nr of waiting loops: %.1f\n", (float)cnt/MILLION); texit = true; pthread_join(tid, NULL); return 0; }
测试结果
nloops=100时
run_x: 0.3049 s run_y: 0.2559 s run_xy: 0.5074 s run_xy_parallel: 0.4744 s average nr of waiting loops: 34.9
nloops=1000时
run_x: 2.4921 s run_y: 2.4965 s run_xy: 5.0950 s run_xy_parallel: 3.3477 s average nr of waiting loops: 44.8
现象原因分析
1. 短任务的同步开销占比过高
当nloops=100时,calc_x和calc_y单次执行时间极短(微秒级),此时自旋同步的开销(atomic_bool读写、自旋等待循环)在总耗时中的占比极高,甚至抵消了并行执行节省的时间,导致几乎看不到收益。
当nloops=1000时,任务单次执行时间变长,同步开销的占比有所降低,所以能看到并行收益,但同步开销依然存在,这就是总耗时比单任务耗时(如run_x的2.49s)慢1/3的核心原因。
2. 原子变量的内存屏障开销
atomic_bool的读写不是普通内存操作,它会触发内存屏障(Memory Barrier),确保多线程间的内存可见性。每次设置trun = true或检查trun时,CPU需要刷新缓存、保证指令顺序,这些操作的耗时固定且不可忽略。百万次循环中每次都要执行这些原子操作,累计开销会随nloops增加被放大,表现为设置trun的耗时上升。
3. 自旋等待的空转损耗
run_xy_parallel中,主线程执行完calc_y后会进入while(trun) ++cnt的自旋等待,CPU一直在空转执行无用的++cnt操作,不仅消耗资源,还可能因为cnt被频繁修改引发缓存竞争,带来额外开销。测试结果中的average nr of waiting loops显示每次等待要循环几十次,这部分空转时间也是总耗时的组成部分。
4. 线程调度的不确定性
工作线程自旋等待trun时,操作系统可能因调度策略暂时切换走该线程,导致主线程完成calc_y后需要等待更长时间才能让工作线程完成calc_x,这种调度延迟在短任务场景下影响尤其明显。
优化建议
- 批量任务处理:不要每次提交单个任务,而是一次性提交一批任务(比如一次提交1000个
calc_x任务),减少原子操作和同步的总次数,降低同步开销占比。 - 使用轻量同步原语:可以用
volatile配合编译器屏障(__sync_synchronize())替代atomic_bool,在保证可见性的前提下减少部分内存屏障开销(需注意正确性)。 - 自旋等待加入pause指令:在自旋循环中加入
__builtin_ia32_pause()(x86平台),减少CPU乱序执行的开销,同时降低功耗,避免被操作系统调度走。 - 采用任务窃取模式:如果有多个工作线程,可使用任务窃取队列,减少主线程和工作线程之间的同步竞争。
内容的提问来源于stack exchange,提问作者user2908112

