Windows线程同步性能异常求助:多线程模拟效率低下
多线程同步问题:Windows与Linux下的性能差异
我开发了一款运行复杂物理模拟的程序,按全年每小时一个工况计算,共需执行8760次模拟。我将这些模拟按线程分组,每个线程平均运行273次模拟循环。
使用AMD Ryzen 9 5950x(16核32线程)执行任务时,Linux系统下所有线程利用率均在98%-100%之间,但Windows系统下出现严重的同步开销问题:

(首条为读取数据的I/O线程,小条为工作线程。红色:同步,绿色:处理,紫色:I/O)
该截图来自Visual Studio并发可视化工具,显示63%的时间花费在线程同步上。Linux和Windows下的代码完全一致,我已尽可能将对象设为不可变,这在旧的8线程Intel i7上带来了显著性能提升,但线程数大幅增加后出现此问题。
多线程实现上,我尝试过自定义ParallelFor以及taskflow库,两者表现完全一致。
是否是Windows线程的本质特性导致了该现象?
自定义ParallelFor代码
/** * parallel for * @tparam Index integer type * @tparam Callable function type * @param start start index of the loop * @param end final +1 index of the loop * @param func function to evaluate * @param nb_threads number of threads, if zero, it is determined automatically */ template<typename Index, typename Callable> static void ParallelFor(Index start, Index end, Callable func, unsigned nb_threads=0) { // Estimate number of threads in the pool if (nb_threads == 0) nb_threads = getThreadNumber(); // Size of a slice for the range functions Index n = end - start + 1; Index slice = (Index) std::round(n / static_cast<double> (nb_threads)); slice = std::max(slice, Index(1)); // [Helper] Inner loop auto launchRange = [&func] (int k1, int k2) { for (Index k = k1; k < k2; k++) { func(k); } }; // Create pool and launch jobs std::vector<std::thread> pool; pool.reserve(nb_threads); Index i1 = start; Index i2 = std::min(start + slice, end); for (unsigned i = 0; i + 1 < nb_threads && i1 < end; ++i) { pool.emplace_back(launchRange, i1, i2); i1 = i2; i2 = std::min(i2 + slice, end); } if (i1 < end) { pool.emplace_back(launchRange, i1, end); } // Wait for jobs to finish for (std::thread &t : pool) { if (t.joinable()) { t.join(); } } }
Main.cpp代码
// // Created by santi on 26/08/2022. // #include "input_data.h" #include "output_data.h" #include "random.h" #include "par_for.h" void fillA(Matrix& A){ Random rnd; rnd.setTimeBasedSeed(); for(int i=0; i < A.getRows(); ++i) for(int j=0; j < A.getRows(); ++j) A(i, j) = (int) rnd.randInt(0, 1000); } void worker(const InputData& input_data, OutputData& output_data, const std::vector<int>& time_indices, int thread_index){ std::cout << "Thread " << thread_index << " [" << time_indices[0]<< ", " << time_indices[time_indices.size() - 1] << "]\n"; for(const int& t: time_indices){ Matrix b = input_data.getAt(t); Matrix A(input_data.getDim(), input_data.getDim()); fillA(A); Matrix x = A * b; output_data.setAt(t, x); } } void process(int time_steps, int dim, int n_threads){ InputData input_data(time_steps, dim); OutputData output_data(time_steps, dim); // correct the number of threads if ( n_threads < 1 ) { n_threads = ( int )getThreadNumber( ); } // generate indices std::vector<int> time_indices = arrange<int>(time_steps); // compute the split of indices per core std::vector<ParallelChunkData<int>> chunks = prepareParallelChunks(time_indices, n_threads ); // run in parallel ParallelFor( 0, ( int )chunks.size( ), [ & ]( int k ) { // run chunk worker(input_data, output_data, chunks[k].indices, k ); } ); } int main(){ process(8760, 5000, 0); return 0; }
问题分析与解决方案
1. 并非Windows线程本质特性,而是同步点实现差异
Linux和Windows的线程调度、同步原语实现细节不同,但核心逻辑一致。问题大概率出在共享资源的访问竞争或线程调度策略适配上,而非系统本身的线程特性。
2. 核心问题定位
从代码和现象来看,同步开销主要来自以下几点:
OutputData::setAt的线程安全问题:多个线程同时写入output_data,如果该方法内部未使用高效同步机制(如无锁结构、分区写入),会导致严重锁竞争。Windows下临界区/互斥锁在高线程数下的开销可能比Linux的pthread锁更明显。std::cout的全局锁:每个线程都调用std::cout输出日志,std::cout是线程安全但带有全局锁的,高线程数下会造成大量同步等待,这在Windows下表现可能更突出。- 线程数量过载:Ryzen 9 5950x是16核32线程(超线程),超线程更适合I/O密集型任务,你的模拟是计算密集型,过多线程会导致核心上下文切换开销激增。Windows调度器在超线程场景下的资源分配策略可能与Linux不同,加剧了同步问题。
3. 具体优化方案
- 移除
std::cout输出:注释掉worker函数中的打印语句,全局锁在高线程数下的开销不可忽视,这是最快速的验证方式。 - 分区写入
OutputData:预先将output_data按线程分配独立存储区域,线程完成任务后再合并,完全避免写入竞争。比如给每个线程分配局部OutputData,最后汇总到全局对象。 - 调整线程数量:尝试将线程数设置为物理核心数(16)而非逻辑线程数(32),计算密集型任务中超线程收益有限,反而会增加调度和同步开销。
- 使用无锁数据结构:如果必须实时写入共享数据,替换
OutputData内部实现为无锁队列或分区数组,避免使用互斥锁。 - 检查
InputData::getAt开销:虽然InputData是不可变的,但如果getAt内部存在隐式同步或缓存失效问题,也会间接导致线程等待。确保getAt是纯读取操作且数据能在缓存中高效访问。
4. 验证步骤
- 注释掉
worker中的std::cout,重新运行测试,查看同步占比是否下降。 - 将线程数手动设置为16,对比32线程的性能差异。
- 实现
OutputData的分区写入,彻底消除写入竞争,观察同步开销是否消失。
内容的提问来源于stack exchange,提问作者Santi Peñate-Vera
相关产品推荐
相关产品推荐

