You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

探究pthreads性能收益阈值:工作负载是否需达毫秒级?

pthread多线程收益阈值:3ms是否合理?

这个3ms左右的阈值是完全合理的,核心原因在于多线程引入的额外开销(线程创建、调度、上下文切换)需要足够的工作负载来分摊,当单线程工作负载低于这个阈值时,额外开销会抵消甚至超过并行带来的收益。

从你的基准测试数据可以清晰看到这个趋势:

  • 单线程运行1.37ms时,双线程版本耗时2.16ms,反而更慢——线程创建和调度的开销完全盖过了并行处理的优势
  • 单线程运行2.75ms时,双线程耗时3.76ms,仍未实现收益
  • 单线程运行4.15ms时,双线程耗时3.67ms,首次实现性能反超,和你观察到的3ms左右阈值基本吻合
  • 当单线程负载达到8.26ms时,双线程耗时4.93ms,接近理想的2倍加速,说明此时工作负载足够大,线程开销占比可以忽略

基准测试输出

BM_dispatch<dispatch>/16/process_time/real_time                1.37 ms         1.37 ms          513
BM_dispatch<dispatch>/32/process_time/real_time                2.75 ms         2.75 ms          252
BM_dispatch<dispatch>/48/process_time/real_time                4.15 ms         4.15 ms          169
BM_dispatch<dispatch>/64/process_time/real_time                5.52 ms         5.52 ms          126
BM_dispatch<dispatch>/80/process_time/real_time                6.89 ms         6.89 ms          101
BM_dispatch<dispatch>/96/process_time/real_time                8.26 ms         8.26 ms           84
BM_dispatch<dispatch>/112/process_time/real_time               9.62 ms         9.62 ms           72

BM_dispatch<dispatch_pthread>/16/process_time/real_time        2.16 ms         4.18 ms          359
BM_dispatch<dispatch_pthread>/32/process_time/real_time        3.76 ms         7.38 ms          200
BM_dispatch<dispatch_pthread>/48/process_time/real_time        3.67 ms         7.18 ms          150
BM_dispatch<dispatch_pthread>/64/process_time/real_time        4.30 ms         8.44 ms          163
BM_dispatch<dispatch_pthread>/80/process_time/real_time        4.38 ms         8.60 ms          176
BM_dispatch<dispatch_pthread>/96/process_time/real_time        4.93 ms         9.69 ms          146
BM_dispatch<dispatch_pthread>/112/process_time/real_time       5.31 ms         10.5 ms          126

测试程序实现

void find_max(const float* in, size_t eles, float* out) {
    float max{0};
    for (size_t i = 0; i < eles; ++i) {
        if (in[i] > max) max = in[i];
    }
    *out = max;
}

float dispatch(const float* inp, size_t rows, size_t cols, float* out) {
    for (size_t row = 0; row < rows; row++) {
        find_max(inp + row * cols, cols, out + row);
    }
}

struct pthreadpool_context {
    const float* inp;
    size_t rows;
    size_t cols;
    float* out;
};

void* work(void* ctx) {
    const pthreadpool_context* context = (pthreadpool_context*)ctx;
    dispatch(context->inp, context->rows, context->cols, context->out);
    return NULL;
}

float dispatch_pthread(const float* inp, size_t rows, size_t cols, float* out) {
    pthread_t thread1, thread2;
    size_t rows_per_thread = rows / 2;
    const pthreadpool_context context1 = {inp, rows_per_thread, cols, out};
    const pthreadpool_context context2 = {inp + rows_per_thread * cols,
                                          rows_per_thread, cols,
                                          out + rows_per_thread};
    pthread_create(&thread1, NULL, work, (void*)&context1);
    pthread_create(&thread2, NULL, work, (void*)&context2);
    pthread_join(thread1, NULL);
    pthread_join(thread2, NULL);
}


template <auto F>
void BM_dispatch(benchmark::State& state) {
    std::random_device rnd_device;
    std::mt19937 mersenne_engine{rnd_device()};
    std::normal_distribution<float> dist{0, 1};
    auto gen = [&]() { return dist(mersenne_engine); };
    const size_t cols = 1024 * state.range(0);
    constexpr size_t rows = 100;
    std::vector<float> inp(rows * cols);
    std::generate(inp.begin(), inp.end(), gen);
    std::vector<float> out(rows);
    for (auto _ : state) {
        F(inp.data(), rows, cols, out.data());
    }
}

BENCHMARK(BM_dispatch<dispatch>)
    ->DenseRange(16, 112, 16)
    ->MeasureProcessCPUTime()
    ->UseRealTime()
    ->Unit(benchmark::kMillisecond);
BENCHMARK(BM_dispatch<dispatch_pthread>)
    ->DenseRange(16, 112, 16)
    ->MeasureProcessCPUTime()
    ->UseRealTime()
    ->Unit(benchmark::kMillisecond);
BENCHMARK_MAIN();

补充说明:你的测试中每次调用dispatch_pthread都会创建新线程,而不是复用线程池,这会进一步拉高收益阈值——pthread_create的开销通常在几微秒到几十微秒级别,但对于毫秒级的小负载来说,这个开销占比非常可观。如果改用线程池复用线程,收益阈值会显著降低(可能降到几百微秒级别)。

内容的提问来源于stack exchange,提问作者fabian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 16:44:52