You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

短任务并行性能异常:POSIX自旋线程耗时疑问

自旋线程并行执行短代码的性能问题分析

问题描述

我需要百万次重复并行执行一段微秒级的短代码,因为创建POSIX线程开销太高,所以采用了自旋工作线程的方案。在简化案例里并行运行calc_x和calc_y,原本期望总耗时接近单任务的耗时,但实际结果没达到预期:当nloops=100时并行几乎没收益,nloops=1000时仍比预期慢1/3,而且设置atomic_bool变量trun的耗时还会随nloops增加而上升,想知道这一现象的原因。

代码示例

#include <pthread.h>
#include <time.h>
#include <stdatomic.h>
#include <stdio.h>

static int nloops = 100;
static int x;
static int y;
static pthread_t tid;
static atomic_bool trun;
static atomic_bool texit;
static int cnt = 0;
static const int MILLION = 1000000;


void calc_x() {
    x = 0;
    for (int i = 0; i < nloops; ++i) {
        x += i;
    }
}

void calc_y() {
    y = 0;
    for (int i = 0; i < nloops; ++i) {
        y += i;
    }
}

void *worker() {
    while (1) {
        if (trun) {
            calc_x();
            trun = false;
        }
        if (texit) {
            break;
        }
    }
    return NULL;
}

float run_x() {
    struct timespec start, end;
    clock_gettime(CLOCK_MONOTONIC, &start);
    for (int i = 0; i < MILLION; ++i) {
        calc_x();
    }
    clock_gettime(CLOCK_MONOTONIC, &end);
    return (end.tv_sec - start.tv_sec) + (end.tv_nsec - start.tv_nsec)*0.000000001;
}

float run_y() {
    struct timespec start, end;
    clock_gettime(CLOCK_MONOTONIC, &start);
    for (int i = 0; i < MILLION; ++i) {
        calc_y();
    }
    clock_gettime(CLOCK_MONOTONIC, &end);
    return (end.tv_sec - start.tv_sec) + (end.tv_nsec - start.tv_nsec)*0.000000001;
}

float run_xy() {
    struct timespec start, end;
    clock_gettime(CLOCK_MONOTONIC, &start);
    for (int i = 0; i < MILLION; ++i) {
        calc_x();
        calc_y();
    }
    clock_gettime(CLOCK_MONOTONIC, &end);
    return (end.tv_sec - start.tv_sec) + (end.tv_nsec - start.tv_nsec)*0.000000001;
}

float run_xy_parallel() {
    struct timespec start, end;
    clock_gettime(CLOCK_MONOTONIC, &start);
    for (int i = 0; i < MILLION; ++i) {
        trun = true;
        calc_y();
        while(trun) ++cnt;
    }
    clock_gettime(CLOCK_MONOTONIC, &end);
    return (end.tv_sec - start.tv_sec) + (end.tv_nsec - start.tv_nsec)*0.000000001;
}

int main() {
    trun = false;
    texit = false;
    pthread_create(&tid, NULL, &worker, NULL);
    printf("run_x: %.4f s\n", run_x());
    printf("run_y: %.4f s\n", run_y());
    printf("run_xy: %.4f s\n", run_xy());
    printf("run_xy_parallel: %.4f s\n", run_xy_parallel());
    printf("average nr of waiting loops: %.1f\n", (float)cnt/MILLION);
    texit = true;
    pthread_join(tid, NULL);
    return 0;
}

测试结果

nloops=100时

run_x: 0.3049 s
run_y: 0.2559 s
run_xy: 0.5074 s
run_xy_parallel: 0.4744 s
average nr of waiting loops: 34.9

nloops=1000时

run_x: 2.4921 s
run_y: 2.4965 s
run_xy: 5.0950 s
run_xy_parallel: 3.3477 s
average nr of waiting loops: 44.8

现象原因分析

1. 短任务的同步开销占比过高

当nloops=100时,calc_x和calc_y单次执行时间极短(微秒级),此时自旋同步的开销(atomic_bool读写、自旋等待循环)在总耗时中的占比极高,甚至抵消了并行执行节省的时间,导致几乎看不到收益。

当nloops=1000时,任务单次执行时间变长,同步开销的占比有所降低,所以能看到并行收益,但同步开销依然存在,这就是总耗时比单任务耗时(如run_x的2.49s)慢1/3的核心原因。

2. 原子变量的内存屏障开销

atomic_bool的读写不是普通内存操作,它会触发内存屏障(Memory Barrier),确保多线程间的内存可见性。每次设置trun = true或检查trun时,CPU需要刷新缓存、保证指令顺序,这些操作的耗时固定且不可忽略。百万次循环中每次都要执行这些原子操作,累计开销会随nloops增加被放大,表现为设置trun的耗时上升。

3. 自旋等待的空转损耗

run_xy_parallel中,主线程执行完calc_y后会进入while(trun) ++cnt的自旋等待,CPU一直在空转执行无用的++cnt操作,不仅消耗资源,还可能因为cnt被频繁修改引发缓存竞争,带来额外开销。测试结果中的average nr of waiting loops显示每次等待要循环几十次,这部分空转时间也是总耗时的组成部分。

4. 线程调度的不确定性

工作线程自旋等待trun时,操作系统可能因调度策略暂时切换走该线程,导致主线程完成calc_y后需要等待更长时间才能让工作线程完成calc_x,这种调度延迟在短任务场景下影响尤其明显。

优化建议

  • 批量任务处理:不要每次提交单个任务,而是一次性提交一批任务(比如一次提交1000个calc_x任务),减少原子操作和同步的总次数,降低同步开销占比。
  • 使用轻量同步原语:可以用volatile配合编译器屏障(__sync_synchronize())替代atomic_bool,在保证可见性的前提下减少部分内存屏障开销(需注意正确性)。
  • 自旋等待加入pause指令:在自旋循环中加入__builtin_ia32_pause()(x86平台),减少CPU乱序执行的开销,同时降低功耗,避免被操作系统调度走。
  • 采用任务窃取模式:如果有多个工作线程,可使用任务窃取队列,减少主线程和工作线程之间的同步竞争。

内容的提问来源于stack exchange,提问作者user2908112

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 23:35:19