You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Windows线程同步性能异常求助:多线程模拟效率低下

多线程同步问题:Windows与Linux下的性能差异

我开发了一款运行复杂物理模拟的程序,按全年每小时一个工况计算,共需执行8760次模拟。我将这些模拟按线程分组,每个线程平均运行273次模拟循环。

使用AMD Ryzen 9 5950x(16核32线程)执行任务时,Linux系统下所有线程利用率均在98%-100%之间,但Windows系统下出现严重的同步开销问题:

Windows线程同步开销截图
(首条为读取数据的I/O线程,小条为工作线程。红色:同步,绿色:处理,紫色:I/O)

该截图来自Visual Studio并发可视化工具,显示63%的时间花费在线程同步上。Linux和Windows下的代码完全一致,我已尽可能将对象设为不可变,这在旧的8线程Intel i7上带来了显著性能提升,但线程数大幅增加后出现此问题。

多线程实现上,我尝试过自定义ParallelFor以及taskflow库,两者表现完全一致。

是否是Windows线程的本质特性导致了该现象?

自定义ParallelFor代码

/**
 * parallel for
 * @tparam Index integer type
 * @tparam Callable function type
 * @param start start index of the loop
 * @param end final +1 index of the loop
 * @param func function to evaluate
 * @param nb_threads number of threads, if zero, it is determined automatically
 */
template<typename Index, typename Callable>
static void ParallelFor(Index start, Index end, Callable func, unsigned nb_threads=0) {

    // Estimate number of threads in the pool
    if (nb_threads == 0) nb_threads = getThreadNumber();

    // Size of a slice for the range functions
    Index n = end - start + 1;
    Index slice = (Index) std::round(n / static_cast<double> (nb_threads));
    slice = std::max(slice, Index(1));

    // [Helper] Inner loop
    auto launchRange = [&func] (int k1, int k2) {
        for (Index k = k1; k < k2; k++) {
            func(k);
        }
    };

    // Create pool and launch jobs
    std::vector<std::thread> pool;
    pool.reserve(nb_threads);
    Index i1 = start;
    Index i2 = std::min(start + slice, end);

    for (unsigned i = 0; i + 1 < nb_threads && i1 < end; ++i) {
        pool.emplace_back(launchRange, i1, i2);
        i1 = i2;
        i2 = std::min(i2 + slice, end);
    }

    if (i1 < end) {
        pool.emplace_back(launchRange, i1, end);
    }

    // Wait for jobs to finish
    for (std::thread &t : pool) {
        if (t.joinable()) {
            t.join();
        }
    }
}

Main.cpp代码

//
// Created by santi on 26/08/2022.
//
#include "input_data.h"
#include "output_data.h"
#include "random.h"
#include "par_for.h"

void fillA(Matrix& A){

    Random rnd;
    rnd.setTimeBasedSeed();

    for(int i=0; i < A.getRows(); ++i)
        for(int j=0; j < A.getRows(); ++j)
            A(i, j) = (int) rnd.randInt(0, 1000);

}


void worker(const InputData& input_data,
            OutputData& output_data,
            const std::vector<int>& time_indices,
            int thread_index){

    std::cout << "Thread " << thread_index << " [" << time_indices[0]<< ", " << time_indices[time_indices.size() - 1] << "]\n";


    for(const int& t: time_indices){

        Matrix b = input_data.getAt(t);

        Matrix A(input_data.getDim(), input_data.getDim());
        fillA(A);

        Matrix x = A * b;

        output_data.setAt(t, x);
    }

}


void process(int time_steps, int dim, int n_threads){
    InputData input_data(time_steps, dim);
    OutputData output_data(time_steps, dim);

    // correct the number of threads
    if ( n_threads < 1 ) { n_threads = ( int )getThreadNumber( ); }

    // generate indices
    std::vector<int> time_indices = arrange<int>(time_steps);

    // compute the split of indices per core
    std::vector<ParallelChunkData<int>> chunks = prepareParallelChunks(time_indices, n_threads );

    // run in parallel
    ParallelFor( 0, ( int )chunks.size( ), [ & ]( int k ) {
            // run chunk
            worker(input_data, output_data, chunks[k].indices, k );
    } );
}

int main(){

    process(8760, 5000, 0);

    return 0;
}

问题分析与解决方案

1. 并非Windows线程本质特性,而是同步点实现差异

Linux和Windows的线程调度、同步原语实现细节不同,但核心逻辑一致。问题大概率出在共享资源的访问竞争或线程调度策略适配上,而非系统本身的线程特性。

2. 核心问题定位

从代码和现象来看,同步开销主要来自以下几点:

  • OutputData::setAt的线程安全问题:多个线程同时写入output_data,如果该方法内部未使用高效同步机制(如无锁结构、分区写入),会导致严重锁竞争。Windows下临界区/互斥锁在高线程数下的开销可能比Linux的pthread锁更明显。
  • std::cout的全局锁:每个线程都调用std::cout输出日志,std::cout是线程安全但带有全局锁的,高线程数下会造成大量同步等待,这在Windows下表现可能更突出。
  • 线程数量过载:Ryzen 9 5950x是16核32线程(超线程),超线程更适合I/O密集型任务,你的模拟是计算密集型,过多线程会导致核心上下文切换开销激增。Windows调度器在超线程场景下的资源分配策略可能与Linux不同,加剧了同步问题。

3. 具体优化方案

  • 移除std::cout输出:注释掉worker函数中的打印语句,全局锁在高线程数下的开销不可忽视,这是最快速的验证方式。
  • 分区写入OutputData:预先将output_data按线程分配独立存储区域,线程完成任务后再合并,完全避免写入竞争。比如给每个线程分配局部OutputData,最后汇总到全局对象。
  • 调整线程数量:尝试将线程数设置为物理核心数(16)而非逻辑线程数(32),计算密集型任务中超线程收益有限,反而会增加调度和同步开销。
  • 使用无锁数据结构:如果必须实时写入共享数据,替换OutputData内部实现为无锁队列或分区数组,避免使用互斥锁。
  • 检查InputData::getAt开销:虽然InputData是不可变的,但如果getAt内部存在隐式同步或缓存失效问题,也会间接导致线程等待。确保getAt是纯读取操作且数据能在缓存中高效访问。

4. 验证步骤

  1. 注释掉worker中的std::cout,重新运行测试,查看同步占比是否下降。
  2. 将线程数手动设置为16,对比32线程的性能差异。
  3. 实现OutputData的分区写入,彻底消除写入竞争,观察同步开销是否消失。

内容的提问来源于stack exchange,提问作者Santi Peñate-Vera

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 01:24:31