16核Intel i9-12900HX向量内积并行计算性能停滞问题咨询
向量内积并行计算性能瓶颈问题
我的设备搭载16核Intel i9-12900HX处理器,为充分利用多核性能,开展了向量内积计算实验:计算两个含10^9个double类型元素的一维向量的内积,分别采用串行单循环、1至8线程并行两种方式。
实验结果
N Threads ----- Plel DotProd --- Plel Time(ms) -- Series DotProd - Series Time(ms) 1 166666666.67 714.55 166666666.67 694.82 2 166666666.67 416.97 3 166666666.67 387.20 4 166666666.67 316.38 5 166666666.67 315.09 6 166666666.67 293.89 7 166666666.67 286.43 8 166666666.67 277.68
结果显示所有内积计算结果一致:
- 串行计算耗时694.82ms,1线程并行耗时714.55ms;
- 1到2线程提速合理,但2到3、3到4线程提速未达预期;
- 线程数≥4后性能无明显提升。
疑问
- 是否所有线程仅占用3-4个核心?
- 如何知晓并确保计算使用的核心数?
- 问题是否出在其他环节?
实验代码
#include <vector> #include <thread> #include <chrono> #include <iostream> #include <iomanip> using std::chrono::high_resolution_clock; using std::chrono::duration; //The functor being parallelized void innerProduct(const int threadID, const long start, const long end, const std::vector<double>& a, const std::vector<double>& b, std::vector<double>& innerProductResult //ith element holds result from ith thread ) { innerProductResult[threadID] = 0.0; for (long i = start; i < end; i++) innerProductResult[threadID] += a[i] * b[i]; } int main(void) { //Create and populate const int DIMENSION = 1'000'000'000; std::vector<double> a(DIMENSION); std::vector<double> b(DIMENSION); //Use values that won't explode the result or invoke any multiplication optimzations for (int i = 0; i < DIMENSION; i++) { double arg = i / (double)DIMENSION; a[i] = arg; b[i] = 1.0 - arg; } //Compute inner product in SERIES double seriesInnerProduct = 0.0; auto t1 = high_resolution_clock::now(); for (int i = 0; i < DIMENSION; i++) seriesInnerProduct += a[i] * b[i]; auto t2 = high_resolution_clock::now(); duration<double, std::milli> millisecs_duration = t2 - t1; //milliseconds as double const double seriesTime = millisecs_duration.count(); //Compute inner product in PARALLEL const int MAX_NUMBER_OF_THREADS = 8; std::vector<double> parallelTimings(MAX_NUMBER_OF_THREADS); std::vector<double> parallelInnerProducts(MAX_NUMBER_OF_THREADS, 0.0); //initialize elements to zero //Compute the inner product using different numbers of threads for (int numThreadsUsed = 1; numThreadsUsed <= MAX_NUMBER_OF_THREADS; numThreadsUsed++) { //Create batches for subjobs const long batch_size = DIMENSION / numThreadsUsed; //but there may be a remainder, so... const long sizeOfLastBatch = batch_size + DIMENSION % numThreadsUsed; //include any remainder std::vector<double> innerProductByThread(numThreadsUsed); std::vector< std::thread > threads(numThreadsUsed); t1 = high_resolution_clock::now(); //Spread the work over the number of threads being used for (int i = 0; i < numThreadsUsed; i++) { long start = i * batch_size; long end = start + (i < numThreadsUsed - 1 ? batch_size : sizeOfLastBatch); threads[i] = std::thread(innerProduct, i, start, end, std::ref(a), std::ref(b), std::ref(innerProductByThread)); } // Wait for everyone to finish for (int i = 0; i < numThreadsUsed; i++) threads[i].join(); //Aggregate the (sub) inner product results from the threads used for (int i = 0; i < numThreadsUsed; i++) parallelInnerProducts[numThreadsUsed - 1] += innerProductByThread[i]; t2 = high_resolution_clock::now(); millisecs_duration = t2 - t1; parallelTimings[numThreadsUsed - 1] = millisecs_duration.count(); } //Output to screen std::cout << " Num Plel Plel Series Series " << std::endl; std::cout << " Threads DotProd Time(ms) DotProd Time(ms)" << std::endl << std::endl; for (int numThreadsUsed = 1; numThreadsUsed <= MAX_NUMBER_OF_THREADS; numThreadsUsed++) { std::cout << std::setw(5) << numThreadsUsed << std::fixed << std::setw(14) << std::setprecision(2) << parallelInnerProducts[numThreadsUsed - 1] << std::fixed << std::setw(9) << std::setprecision(2) << parallelTimings[numThreadsUsed - 1]; if (numThreadsUsed == 1) std::cout << std::fixed << std::setw(13) << std::setprecision(2) << seriesInnerProduct << std::setw(8) << std::setprecision(2) << seriesTime; std::cout << std::endl; } std::cout << std::endl << "Press Enter to exit..."; std::cin.get(); return 1; }
问题分析与解决方案
1. 线程是否仅占用3-4个核心?
不是,核心利用率不是瓶颈,你的实验受限于内存带宽饱和。i9-12900HX的内存带宽有限(DDR5-4800约76.8GB/s),而向量内积是典型的内存密集型任务:10^9次循环需要读取16GB数据、写入8GB数据,串行计算时内存带宽已经接近上限,增加线程数无法获取更多带宽,性能自然不再提升。另外1线程并行比串行慢,是因为线程创建、调度有额外开销,且串行代码可能被编译器优化得更彻底(比如自动向量化)。
2. 如何查看并确保线程使用的核心数?
- Windows:打开任务管理器→详细信息,找到目标进程,右键→设置相关性,可查看/修改线程绑定的核心;也可用命令
wmic process get processid,threadcount,affinity查询。 - Linux/macOS:用
htop(按H显示线程)、ps -eLf | grep <进程名>查看线程的CPU亲和性;或用taskset命令查看/设置核心绑定。
要确保计算使用更多核心,可手动设置线程亲和性:Windows用SetThreadAffinityMask,Linux用pthread_setaffinity_np,把每个线程绑定到不同的物理核心(优先绑定i9-12900HX的8个性能核,而非能效核);也可以改用OpenMP、Intel TBB这类并行库,它们会自动管理线程亲和性与负载均衡,比手动创建线程更高效。
3. 其他可能的问题点
- 编译器优化:确保编译时开启最高优化级别(如
-O3或/O2),自动向量化能大幅提升单线程性能,为并行计算打下更好基础。 - 大小核调度:i9-12900HX是大小核架构,系统默认调度可能把线程放到能效核,导致性能受限。可在任务管理器中将进程设为“高优先级”,或手动绑定到性能核。
- 缓存利用率:你的向量是连续内存,缓存命中率较高,但如果线程划分的块过小可能引发缓存行竞争,不过你的batch size(125MB)远大于L3缓存(24MB),这一影响可以忽略。
内容的提问来源于stack exchange,提问作者RickB88
相关产品推荐
相关产品推荐

