C++多线程计算时间递增问题的示例分析与原因问询
问题
我在处理大规模图实例的复杂问题时,采用多线程拆分输入空间,让各线程独立执行同一函数。测试软件扩展性时发现,线程数超过4后,计算时间随线程数增加而上升。
为排查原因,我编写了如下C++示例代码:
#include <algorithm> #include <random> #include <thread> #include <iostream> #include <chrono> template<typename T> inline double getMs(T start, T end) { return double( std::chrono::duration_cast<std::chrono::milliseconds>(end - start) .count()) / 1000; } int main(int) { std::random_device rd; std::mt19937 g(rd()); unsigned int n = std::thread::hardware_concurrency(); std::cout << n << " concurrent threads are supported.\n"; for (size_t np = 2; np < 17; np++) { auto start = std::chrono::high_resolution_clock::now(); std::cout << np << " threads: "; std::vector<std::thread> threads(np); int number_stops = 50; // memory 39420 int number_transfers = 1; // memory int number_structures = 1; // memory int number_iterations = 1000000; // time auto dimension = number_stops * (number_transfers + 1) * number_structures; auto paraTask = [&]() { for (int b = 0; b < number_iterations; b++) { //std::srand(unsigned(std::time(nullptr))); std::vector<int> v(dimension, 1586); //std::generate(v.begin(), v.end(), std::rand); v.clear(); } }; for (size_t i = 0; i < np; i++) { threads[i] = std::thread(paraTask); } // Join the threads for (auto&& thread : threads) thread.join(); double elapsed = getMs(start, std::chrono::high_resolution_clock::now()); printf("parallel completed: %.3f sec.\n", elapsed); } return 0; }
示例中,number_stops、number_transfers、number_structures用于模拟内存消耗,number_iterations模拟计算迭代次数。两种参数配置下的测试结果均显示线程数超过4后计算时间递增:
第一种配置测试结果
16 concurrent threads are supported. 2 threads: parallel completed: 0.995 sec. 3 threads: parallel completed: 1.017 sec. 4 threads: parallel completed: 1.028 sec. 5 threads: parallel completed: 1.081 sec. 6 threads: parallel completed: 1.131 sec. 7 threads: parallel completed: 1.122 sec. 8 threads: parallel completed: 1.216 sec. 9 threads: parallel completed: 1.445 sec. 10 threads: parallel completed: 1.603 sec. 11 threads: parallel completed: 1.596 sec. 12 threads: parallel completed: 1.626 sec. 13 threads: parallel completed: 1.634 sec. 14 threads: parallel completed: 1.611 sec. 15 threads: parallel completed: 1.648 sec. 16 threads: parallel completed: 1.688 sec.
第二种配置(高内存低迭代)测试结果
16 concurrent threads are supported. 2 threads: parallel completed: 0.275 sec. 3 threads: parallel completed: 0.267 sec. 4 threads: parallel completed: 0.278 sec. 5 threads: parallel completed: 0.282 sec. 6 threads: parallel completed: 0.303 sec. 7 threads: parallel completed: 0.314 sec. 8 threads: parallel completed: 0.345 sec. 9 threads: parallel completed: 0.370 sec. 10 threads: parallel completed: 0.368 sec. 11 threads: parallel completed: 0.395 sec. 12 threads: parallel completed: 0.407 sec. 13 threads: parallel completed: 0.431 sec. 14 threads: parallel completed: 0.444 sec. 15 threads: parallel completed: 0.448 sec. 16 threads: parallel completed: 0.455 sec.
硬件配置
- CPU:11th Gen Intel(R) Core(TM) i7-11700KF @ 3.60GHz(8物理核心,16逻辑核心)
- RAM:16 GB DDR4
- 系统与编译器:Windows 11,MS_VS 2022
请问为何会出现线程数增加后计算时间递增的情况?
原因分析
- 内存带宽瓶颈:测试代码核心是频繁创建、销毁vector,本质是大量内存分配与释放操作。DDR4内存带宽有限,当线程数超过4后,多线程同时争抢内存带宽,导致内存访问延迟上升,整体耗时增加。尤其是高内存配置的测试中,内存竞争的影响更显著。
- 超线程的局限性:你的CPU有8个物理核心、16个逻辑核心(超线程)。超线程让单个物理核心同时处理两个线程,但这两个线程共享物理核心的执行资源(如缓存、运算单元)。当线程数超过物理核心数的一半(4)后,新增线程开始占用超线程资源,物理核心无法完全并行支撑多线程执行,反而会因资源争抢、线程切换带来额外开销。
- 线程调度开销:线程数越多,操作系统调度压力越大,需要频繁切换线程上下文。每个上下文切换都要保存和恢复寄存器、内存状态,消耗CPU时间。当线程数超过4后,这种调度开销开始凸显,抵消甚至超过多线程并行的收益。
- 缓存命中率下降:每个物理核心的L1/L2缓存是独占的,L3缓存是共享的。线程数增多时,更多线程占用缓存空间,导致单个线程的缓存命中率下降,更多内存访问需要从主存获取,而主存访问速度远低于缓存,进而增加整体耗时。
内容的提问来源于stack exchange,提问作者Claudio Tomasi
相关产品推荐
相关产品推荐

