探究pthreads性能收益阈值:工作负载是否需达毫秒级?
pthread多线程收益阈值:3ms是否合理?
这个3ms左右的阈值是完全合理的,核心原因在于多线程引入的额外开销(线程创建、调度、上下文切换)需要足够的工作负载来分摊,当单线程工作负载低于这个阈值时,额外开销会抵消甚至超过并行带来的收益。
从你的基准测试数据可以清晰看到这个趋势:
- 单线程运行1.37ms时,双线程版本耗时2.16ms,反而更慢——线程创建和调度的开销完全盖过了并行处理的优势
- 单线程运行2.75ms时,双线程耗时3.76ms,仍未实现收益
- 单线程运行4.15ms时,双线程耗时3.67ms,首次实现性能反超,和你观察到的3ms左右阈值基本吻合
- 当单线程负载达到8.26ms时,双线程耗时4.93ms,接近理想的2倍加速,说明此时工作负载足够大,线程开销占比可以忽略
基准测试输出
BM_dispatch<dispatch>/16/process_time/real_time 1.37 ms 1.37 ms 513 BM_dispatch<dispatch>/32/process_time/real_time 2.75 ms 2.75 ms 252 BM_dispatch<dispatch>/48/process_time/real_time 4.15 ms 4.15 ms 169 BM_dispatch<dispatch>/64/process_time/real_time 5.52 ms 5.52 ms 126 BM_dispatch<dispatch>/80/process_time/real_time 6.89 ms 6.89 ms 101 BM_dispatch<dispatch>/96/process_time/real_time 8.26 ms 8.26 ms 84 BM_dispatch<dispatch>/112/process_time/real_time 9.62 ms 9.62 ms 72 BM_dispatch<dispatch_pthread>/16/process_time/real_time 2.16 ms 4.18 ms 359 BM_dispatch<dispatch_pthread>/32/process_time/real_time 3.76 ms 7.38 ms 200 BM_dispatch<dispatch_pthread>/48/process_time/real_time 3.67 ms 7.18 ms 150 BM_dispatch<dispatch_pthread>/64/process_time/real_time 4.30 ms 8.44 ms 163 BM_dispatch<dispatch_pthread>/80/process_time/real_time 4.38 ms 8.60 ms 176 BM_dispatch<dispatch_pthread>/96/process_time/real_time 4.93 ms 9.69 ms 146 BM_dispatch<dispatch_pthread>/112/process_time/real_time 5.31 ms 10.5 ms 126
测试程序实现
void find_max(const float* in, size_t eles, float* out) { float max{0}; for (size_t i = 0; i < eles; ++i) { if (in[i] > max) max = in[i]; } *out = max; } float dispatch(const float* inp, size_t rows, size_t cols, float* out) { for (size_t row = 0; row < rows; row++) { find_max(inp + row * cols, cols, out + row); } } struct pthreadpool_context { const float* inp; size_t rows; size_t cols; float* out; }; void* work(void* ctx) { const pthreadpool_context* context = (pthreadpool_context*)ctx; dispatch(context->inp, context->rows, context->cols, context->out); return NULL; } float dispatch_pthread(const float* inp, size_t rows, size_t cols, float* out) { pthread_t thread1, thread2; size_t rows_per_thread = rows / 2; const pthreadpool_context context1 = {inp, rows_per_thread, cols, out}; const pthreadpool_context context2 = {inp + rows_per_thread * cols, rows_per_thread, cols, out + rows_per_thread}; pthread_create(&thread1, NULL, work, (void*)&context1); pthread_create(&thread2, NULL, work, (void*)&context2); pthread_join(thread1, NULL); pthread_join(thread2, NULL); } template <auto F> void BM_dispatch(benchmark::State& state) { std::random_device rnd_device; std::mt19937 mersenne_engine{rnd_device()}; std::normal_distribution<float> dist{0, 1}; auto gen = [&]() { return dist(mersenne_engine); }; const size_t cols = 1024 * state.range(0); constexpr size_t rows = 100; std::vector<float> inp(rows * cols); std::generate(inp.begin(), inp.end(), gen); std::vector<float> out(rows); for (auto _ : state) { F(inp.data(), rows, cols, out.data()); } } BENCHMARK(BM_dispatch<dispatch>) ->DenseRange(16, 112, 16) ->MeasureProcessCPUTime() ->UseRealTime() ->Unit(benchmark::kMillisecond); BENCHMARK(BM_dispatch<dispatch_pthread>) ->DenseRange(16, 112, 16) ->MeasureProcessCPUTime() ->UseRealTime() ->Unit(benchmark::kMillisecond); BENCHMARK_MAIN();
补充说明:你的测试中每次调用dispatch_pthread都会创建新线程,而不是复用线程池,这会进一步拉高收益阈值——pthread_create的开销通常在几微秒到几十微秒级别,但对于毫秒级的小负载来说,这个开销占比非常可观。如果改用线程池复用线程,收益阈值会显著降低(可能降到几百微秒级别)。
内容的提问来源于stack exchange,提问作者fabian
相关产品推荐
相关产品推荐

