为何核协方差矩阵多线程计算未比单核提速?
多线程加速核协方差矩阵计算无性能提升问题
这是我首次尝试通过多线程加速重型计算。
背景:需基于3D点列表x_test计算核协方差矩阵(Kernel Covariance matrix),矩阵维度为x_test.size()×x_test.size()。
我已通过仅计算下三角矩阵(lower triangular matrix)实现提速,因所有计算相互独立,尝试按行拆分矩阵计算任务分配给多线程(当前x_test.size() = 27000)以进一步提速。
单核计算耗时约280秒,4核计算耗时270-290秒
main.cpp
int main(int argc, char *argv[]) { double sigma0sq = 1; double lengthScale [] = {0.7633, 0.6937, 3.3307e+07}; const std::vector<std::vector<double>> x_test = parse2DCsvFile(inputPath); /* Finding data slices of similar size */ //This piece of code works, each thread is assigned roughly the same number of matrix entries int numElements = x_test.size()*x_test.size()/2; const int numThreads = 4; int elemsPerThread = numElements / numThreads; std::vector<int> indices; int j = 0; for(std::size_t i=1; i<x_test.size()+1; ++i){ int prod = i*(i+1)/2 - j*(j+1)/2; if (prod > elemsPerThread) { i--; j = i; indices.push_back(i); if(indices.size() == numThreads-1) break; } } indices.insert(indices.begin(), 0); indices.push_back(x_test.size()); /* Spreding calculations to multiple threads */ std::vector<std::thread> threads; for(std::size_t i = 1; i < indices.size(); ++i){ threads.push_back(std::thread(calculateKMatrixCpp, x_test, lengthScale, sigma0sq, i, indices.at(i-1), indices.at(i))); } for(auto & th: threads){ th.join(); } return 0; }
每个线程对分配的数据执行以下计算:
calculateKMatrixCpp 函数
void calculateKMatrixCpp(const std::vector<std::vector<double>> xtest, double lengthScale[], double sigma0sq, int threadCounter, int start, int stop){ char buffer[8192]; std::ofstream out("lower_half_matrix_" + std::to_string(threadCounter) +".csv"); out.rdbuf()->pubsetbuf(buffer, 8196); for(int i = start; i < stop; ++i){ for(int j = 0; j < i+1; ++j){ double kij = seKernel(xtest.at(i), xtest.at(j), lengthScale, sigma0sq); if (j!=0) out << ','; out << kij; } if(i!=xtest.size()-1 ) out << '\n'; } out.close(); }
seKernel 函数
double seKernel(const std::vector<double> x1,const std::vector<double> x2, double lengthScale[], double sigma0sq) { double sum(0); for(std::size_t i=0; i<x1.size();i++){ sum += pow((x1.at(i)-x2.at(i))/lengthScale[i],2); } return sigma0sq*exp(-0.5*sum); }
已考虑因素
- 数据向量并发访问锁问题:未向线程传递引用,而是传递数据副本,虽内存占用非最优,但可避免并发数据访问
- 输出:每个线程将下三角矩阵的对应部分写入独立文件,任务管理器显示SSD未达满负荷
编译器与设备环境
- Windows 11
- GNU GCC Compiler
- Code::Blocks(认为影响不大)
内容的提问来源于stack exchange,提问作者blockchain187
相关产品推荐
相关产品推荐

