基于OpenMP的代码为何无性能提升?求技术分析
问题描述
我尝试将“有效点”填充到vector<vector<int>>中,采用先使用本地vector<vector<int>>再合并的方式,附上完整代码和测试结果后,发现双线程理论上读取缓冲区的工作量减半,但最终总耗时几乎与单线程相同,求原因及解决方法。
测试代码
const char* pcData = pkt.data().c_str(); // 大字符串缓冲区 int numChannels = 128; int numThreads = 1; std::vector<std::vector<int>> rowGlobalIds(numChannels, std::vector<int>()); for(auto& v: rowGlobalIds){ v.reserve(3000); } #pragma omp parallel num_threads(numThreads) { omp_set_nested(1); double wtime = omp_get_wtime(); int thread_num = omp_get_thread_num(); int localItNumber = 0; int localValidNumber = 0; // 创建本地vector,后续合并到全局容器 std::vector<std::vector<int>> thIds(numChannels, std::vector<int>()); for(auto& v: thIds){ v.reserve(1500); } #pragma omp for nowait schedule(static) for(int i=0; i < pointCount; ++i){ int16_t val1 = copyFromBuffer(pcData, i, 1); if(val1 > threshold){ uint16_t row = copyFromBuffer(pcData, i, 2); thIds[row].push_back(i); localValidNumber++; } localItNumber++; } wtime = omp_get_wtime() - wtime; printf( "Time taken by thread %d is %f ms with %d iterations and %d valid points\n", thread_num, wtime*1000, itcounts[thread_num], validPoints[thread_num] ); // 合并本地vector到全局容器,要求按线程0、1的顺序执行 #pragma omp for schedule(static) ordered for(int th = 0; th < omp_get_num_threads(); ++th) { for(int i = 0; i < numChannels; ++i){ #pragma omp ordered rowGlobalIds[i].insert(rowGlobalIds[i].end(), thIds[i].begin(), thIds[i].end()); } } omp_set_nested(0); }
测试结果
当缓冲区包含约600,000个数据位时:
numThreads = 1时:
- Thread 0:
- 迭代次数:634,060
- 有效点数:214,742
- 运行时间:33.936843 ms
numThreads = 2时:
- Thread 0:
- 迭代次数:317,644
- 有效点数:112,428
- 运行时间:34.031533 ms
- Thread 1:
- 迭代次数:317,644
- 有效点数:102,304
- 运行时间:26.162064 ms
原因分析
1. 合并阶段逻辑错误且完全串行化
你的合并代码存在严重逻辑问题:
- 每个线程的
thIds是本地变量,但合并循环让所有线程遍历th(线程编号),并使用**当前线程自己的thIds**去合并,这会导致数据重复插入或缺失,完全不符合预期。 - 即使逻辑修正,
#pragma omp ordered强制所有合并操作按线程编号顺序串行执行,这意味着双线程下合并阶段是完全串行的,抵消了并行循环带来的性能收益。而且线程0会等待整个合并过程完成,所以它的总耗时几乎和单线程一致。
2. 计时范围不准确
线程的计时包含了后续的合并阶段开销,而双线程下线程0需要等待串行合并完成,导致其计时几乎等于单线程总耗时;线程1虽先完成并行循环,但后续也会参与合并等待,最终程序总耗时由线程0的耗时决定。
3. 内存访问模式可能限制性能提升
copyFromBuffer如果是按i偏移读取非连续内存,会导致缓存命中率较低。单线程和双线程下缓存命中率差异不大,因此并行带来的性能提升被内存访问瓶颈抵消。
优化方案
1. 修正合并逻辑,避免串行化
每个线程只需要合并自己的本地thIds到全局容器,不需要串行执行。为避免多个线程同时写入同一个rowGlobalIds[i]导致数据竞争,可对每个线程的合并操作加临界区保护:
// 替换原合并代码 #pragma omp critical { for(int i = 0; i < numChannels; ++i){ rowGlobalIds[i].insert(rowGlobalIds[i].end(), thIds[i].begin(), thIds[i].end()); } }
如果临界区开销过大,可改为对每个row单独加锁,或使用线程安全的容器(如tbb::concurrent_vector)。
2. 调整计时范围,仅测量并行计算阶段
将计时范围限定在并行循环部分,排除合并阶段的影响,准确评估并行计算的性能:
double wtime = omp_get_wtime(); #pragma omp for nowait schedule(static) for(int i=0; i < pointCount; ++i){ // ... 循环逻辑 ... } wtime = omp_get_wtime() - wtime; // 此时wtime仅为并行循环的耗时
3. 优化内存访问模式
- 确保
copyFromBuffer的内存访问是连续的,可将缓冲区数据按结构化方式存储(如结构体数组),而非通过偏移读取,提升缓存命中率。 - 预分配足够的内存:全局
rowGlobalIds的reserve值可根据实际有效点数调整,避免合并时的内存分配开销。
4. 移除不必要的嵌套并行
代码中omp_set_nested(1)开启了嵌套并行,但当前逻辑不需要嵌套并行,反而会增加线程管理开销,建议移除该设置。
内容的提问来源于stack exchange,提问作者tricostume
相关产品推荐
相关产品推荐

