You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何核协方差矩阵多线程计算未比单核提速?

多线程加速核协方差矩阵计算无性能提升问题

这是我首次尝试通过多线程加速重型计算。
背景:需基于3D点列表x_test计算核协方差矩阵(Kernel Covariance matrix),矩阵维度为x_test.size()×x_test.size()。
我已通过仅计算下三角矩阵(lower triangular matrix)实现提速,因所有计算相互独立,尝试按行拆分矩阵计算任务分配给多线程(当前x_test.size() = 27000)以进一步提速。
单核计算耗时约280秒,4核计算耗时270-290秒

main.cpp

int main(int argc, char *argv[]) {
    double sigma0sq = 1;
    double lengthScale [] = {0.7633, 0.6937, 3.3307e+07};
    
    const std::vector<std::vector<double>> x_test = parse2DCsvFile(inputPath);
    
    /* Finding data slices of similar size */
    //This piece of code works, each thread is assigned roughly the same number of matrix entries
    int numElements = x_test.size()*x_test.size()/2;
    const int numThreads = 4;
    int elemsPerThread = numElements / numThreads;
    std::vector<int> indices;

    int j = 0;
    for(std::size_t i=1; i<x_test.size()+1; ++i){
        int prod = i*(i+1)/2 - j*(j+1)/2;
        if (prod > elemsPerThread) {
            i--;
            j = i;
            indices.push_back(i);
            if(indices.size() == numThreads-1)
                break;
        }
    }
    indices.insert(indices.begin(), 0);
    indices.push_back(x_test.size());

    /* Spreding calculations to multiple threads */
    std::vector<std::thread> threads;

    for(std::size_t i = 1; i < indices.size(); ++i){
        threads.push_back(std::thread(calculateKMatrixCpp, x_test, lengthScale, sigma0sq, i, indices.at(i-1), indices.at(i)));
    }

    for(auto & th: threads){
        th.join();
    }
    return 0;
}

每个线程对分配的数据执行以下计算:

calculateKMatrixCpp 函数

void calculateKMatrixCpp(const std::vector<std::vector<double>> xtest, double lengthScale[], double sigma0sq, int threadCounter, int start, int stop){
    char buffer[8192];

    std::ofstream out("lower_half_matrix_" + std::to_string(threadCounter) +".csv");
    out.rdbuf()->pubsetbuf(buffer, 8196);

    for(int i = start; i < stop; ++i){
        for(int j = 0; j < i+1; ++j){
            double kij = seKernel(xtest.at(i), xtest.at(j), lengthScale, sigma0sq);
            if (j!=0)
                out << ',';
            out << kij;
        }
        if(i!=xtest.size()-1 )
            out << '\n';
    }
    out.close();
}

seKernel 函数

double seKernel(const std::vector<double> x1,const std::vector<double> x2, double lengthScale[], double sigma0sq) {
    double sum(0);
    for(std::size_t i=0; i<x1.size();i++){
        sum += pow((x1.at(i)-x2.at(i))/lengthScale[i],2);
    }
    return sigma0sq*exp(-0.5*sum);
}

已考虑因素

  • 数据向量并发访问锁问题:未向线程传递引用,而是传递数据副本,虽内存占用非最优,但可避免并发数据访问
  • 输出:每个线程将下三角矩阵的对应部分写入独立文件,任务管理器显示SSD未达满负荷

编译器与设备环境

  • Windows 11
  • GNU GCC Compiler
  • Code::Blocks(认为影响不大)

内容的提问来源于stack exchange,提问作者blockchain187

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 07:35:22