You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

16核Intel i9-12900HX向量内积并行计算性能停滞问题咨询

向量内积并行计算性能瓶颈问题

我的设备搭载16核Intel i9-12900HX处理器,为充分利用多核性能,开展了向量内积计算实验:计算两个含10^9个double类型元素的一维向量的内积,分别采用串行单循环、1至8线程并行两种方式。

实验结果

N Threads ----- Plel DotProd --- Plel Time(ms)  -- Series DotProd - Series Time(ms)
   
    1           166666666.67     714.55      166666666.67     694.82
    2           166666666.67     416.97
    3           166666666.67     387.20
    4           166666666.67     316.38
    5           166666666.67     315.09
    6           166666666.67     293.89
    7           166666666.67     286.43
    8           166666666.67     277.68

结果显示所有内积计算结果一致:

  • 串行计算耗时694.82ms,1线程并行耗时714.55ms;
  • 1到2线程提速合理,但2到3、3到4线程提速未达预期;
  • 线程数≥4后性能无明显提升。

疑问

  1. 是否所有线程仅占用3-4个核心?
  2. 如何知晓并确保计算使用的核心数?
  3. 问题是否出在其他环节?

实验代码

#include <vector>
#include <thread>
#include <chrono>
#include <iostream>
#include <iomanip>

using std::chrono::high_resolution_clock;
using std::chrono::duration;

//The functor being parallelized
void innerProduct(const int  threadID,
                  const long start,
                  const long end,
                  const std::vector<double>& a,
                  const std::vector<double>& b,
                  std::vector<double>& innerProductResult //ith element holds result from ith thread
                 )
{
    innerProductResult[threadID] = 0.0;
    for (long i = start; i < end; i++)
        innerProductResult[threadID] += a[i] * b[i];
}


int main(void)
{
    //Create and populate
    const int DIMENSION = 1'000'000'000;
    std::vector<double> a(DIMENSION);
    std::vector<double> b(DIMENSION);

    //Use values that won't explode the result or invoke any multiplication optimzations
    for (int i = 0; i < DIMENSION; i++)
    {
        double arg = i / (double)DIMENSION;
        a[i] = arg;
        b[i] = 1.0 - arg;
    }

    //Compute inner product in SERIES
    double seriesInnerProduct = 0.0;
    auto t1 = high_resolution_clock::now();

    for (int i = 0; i < DIMENSION; i++)
        seriesInnerProduct += a[i] * b[i];

    auto t2 = high_resolution_clock::now();
    duration<double, std::milli> millisecs_duration = t2 - t1;   //milliseconds as double
    const double seriesTime = millisecs_duration.count();

    //Compute inner product in PARALLEL
    const int MAX_NUMBER_OF_THREADS = 8;
    std::vector<double> parallelTimings(MAX_NUMBER_OF_THREADS);
    std::vector<double> parallelInnerProducts(MAX_NUMBER_OF_THREADS, 0.0); //initialize elements to zero

    //Compute the inner product using different numbers of threads
    for (int numThreadsUsed = 1; numThreadsUsed <= MAX_NUMBER_OF_THREADS; numThreadsUsed++)
    {
        //Create batches for subjobs
        const long batch_size = DIMENSION / numThreadsUsed;   //but there may be a remainder, so...
        const long sizeOfLastBatch = batch_size + DIMENSION % numThreadsUsed; //include any remainder

        std::vector<double> innerProductByThread(numThreadsUsed);
        std::vector< std::thread > threads(numThreadsUsed);

        t1 = high_resolution_clock::now();

        //Spread the work over the number of threads being used
        for (int i = 0; i < numThreadsUsed; i++)
        {
            long start = i * batch_size;
            long end = start + (i < numThreadsUsed - 1 ? batch_size : sizeOfLastBatch);

            threads[i] = std::thread(innerProduct, i, start, end, std::ref(a), std::ref(b), std::ref(innerProductByThread));
        }

        // Wait for everyone to finish   
        for (int i = 0; i < numThreadsUsed; i++)
            threads[i].join();

        //Aggregate the (sub) inner product results from the threads used
        for (int i = 0; i < numThreadsUsed; i++)
            parallelInnerProducts[numThreadsUsed - 1] += innerProductByThread[i];

        t2 = high_resolution_clock::now();
        millisecs_duration = t2 - t1;

        parallelTimings[numThreadsUsed - 1] = millisecs_duration.count();
    }

    //Output to screen
    std::cout << "   Num      Plel        Plel     Series     Series " << std::endl;
    std::cout << " Threads  DotProd     Time(ms)   DotProd   Time(ms)" << std::endl << std::endl;

    for (int numThreadsUsed = 1; numThreadsUsed <= MAX_NUMBER_OF_THREADS; numThreadsUsed++)
    {
        std::cout << std::setw(5) << numThreadsUsed
                  << std::fixed << std::setw(14) << std::setprecision(2) << parallelInnerProducts[numThreadsUsed - 1]
                  << std::fixed << std::setw(9) << std::setprecision(2) << parallelTimings[numThreadsUsed - 1];

        if (numThreadsUsed == 1)
            std::cout << std::fixed << std::setw(13) << std::setprecision(2) << seriesInnerProduct
                      << std::setw(8) << std::setprecision(2) << seriesTime;
        
        std::cout << std::endl;

            
    }
  
  
    std::cout << std::endl << "Press Enter to exit...";
    std::cin.get();
    return 1;
}

问题分析与解决方案

1. 线程是否仅占用3-4个核心?

不是,核心利用率不是瓶颈,你的实验受限于内存带宽饱和。i9-12900HX的内存带宽有限(DDR5-4800约76.8GB/s),而向量内积是典型的内存密集型任务:10^9次循环需要读取16GB数据、写入8GB数据,串行计算时内存带宽已经接近上限,增加线程数无法获取更多带宽,性能自然不再提升。另外1线程并行比串行慢,是因为线程创建、调度有额外开销,且串行代码可能被编译器优化得更彻底(比如自动向量化)。

2. 如何查看并确保线程使用的核心数?

  • Windows:打开任务管理器→详细信息,找到目标进程,右键→设置相关性,可查看/修改线程绑定的核心;也可用命令wmic process get processid,threadcount,affinity查询。
  • Linux/macOS:用htop(按H显示线程)、ps -eLf | grep <进程名>查看线程的CPU亲和性;或用taskset命令查看/设置核心绑定。

要确保计算使用更多核心,可手动设置线程亲和性:Windows用SetThreadAffinityMask,Linux用pthread_setaffinity_np,把每个线程绑定到不同的物理核心(优先绑定i9-12900HX的8个性能核,而非能效核);也可以改用OpenMP、Intel TBB这类并行库,它们会自动管理线程亲和性与负载均衡,比手动创建线程更高效。

3. 其他可能的问题点

  • 编译器优化:确保编译时开启最高优化级别(如-O3或/O2),自动向量化能大幅提升单线程性能,为并行计算打下更好基础。
  • 大小核调度:i9-12900HX是大小核架构,系统默认调度可能把线程放到能效核,导致性能受限。可在任务管理器中将进程设为“高优先级”,或手动绑定到性能核。
  • 缓存利用率:你的向量是连续内存,缓存命中率较高,但如果线程划分的块过小可能引发缓存行竞争,不过你的batch size(125MB)远大于L3缓存(24MB),这一影响可以忽略。

内容的提问来源于stack exchange,提问作者RickB88

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 16:09:52