You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

WSL2中C++多线程程序无性能提升的问题排查与解决

WSL2中多线程pthread程序无性能提升的解决办法

问题场景

我在WSL2中运行以下基于pthread的C++多线程程序:

#include <iostream>
#include <pthread.h>
#define N 1024*1024*1024
#define num_threads 8
using namespace std;

long x[num_threads] = {0};

// thread function prototype
void* threadnfunc (void *);
void func (int);

int main() {
    cout << "With threads" << endl;
    pthread_t my_threads[num_threads];
    int arg_to_be_passed_to_func[num_threads];      // idx of array x

    // create threads
    for(int i=0; i<num_threads; i++){
        arg_to_be_passed_to_func[i] = i;
        pthread_create (&my_threads[i], NULL, threadnfunc, &arg_to_be_passed_to_func[i]);
    }
    
    // start threads
    for(int i=0; i<num_threads; i++){
        pthread_join (my_threads[i], NULL);
    }

    // cout << "Without threads" << endl;
    // for(int i=0; i<num_threads; i++){
    //     func(i);
    // }

    return 0;
}

void* threadnfunc (void *arg) {
    int j = *(int*)arg;
    cout << "Starting thread " << j << endl;
    for(int i=0; i<N; i++) x[j]+=1;
    cout << "Thread " << j << " execution completed!" << endl;
    return NULL;
}

void func(int j) {
    cout << "Starting thread " << j << endl;
    for(int i=0; i<N; i++) x[j]+=1;
    cout << "Thread " << j << " execution completed!" << endl;
}

原本预期多线程版本相比单线程会显著提速,但在WSL2中两者执行时间几乎一致,而原生Ubuntu虚拟机上能看到明显加速效果。

系统信息

  • 处理器:Intel Core i7-10750H
  • 基准时钟频率:2.60GHz
  • 内存:16GB
  • 物理核心数:6
  • 逻辑处理器数:12

WSL配置与CPU状态

  • WSL内执行nproc返回12,已通过%USERPROFILE%\.wslconfig指定processors=12,确认WSL可访问所有逻辑处理器
  • BIOS中已开启虚拟化与超线程
  • Windows任务管理器显示所有CPU(0-11)利用率为30-35%
  • cat /proc/cpuinfo | grep MHz显示每个CPU频率为2592MHz,接近基准时钟
  • lscpu输出如下:
Architecture:             x86_64
  CPU op-mode(s):         32-bit, 64-bit
  Address sizes:          39 bits physical, 48 bits virtual
  Byte Order:             Little Endian
CPU(s):                   12
  On-line CPU(s) list:    0-11
Vendor ID:                GenuineIntel
  Model name:             Intel(R) Core(TM) i7-10750H CPU @ 2.60GHz
    CPU family:           6
    Model:                165
    Thread(s) per core:   2
    Core(s) per socket:   6
    Socket(s):            1
    Stepping:             2
    BogoMIPS:             5184.01
    Flags:                fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse s
                          se2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon rep_good nopl xtopology cpuid
                           pni pclmulqdq vmx ssse3 fma cx16 pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f1
                          6c rdrand hypervisor lahf_lm abm 3dnowprefetch invpcid_single ssbd ibrs ibpb stibp ibrs_enhanc
                          ed tpr_shadow vnmi ept vpid ept_ad fsgsbase bmi1 avx2 smep bmi2 erms invpcid rdseed adx smap c
                          lflushopt xsaveopt xsavec xgetbv1 xsaves md_clear flush_l1d arch_capabilities
Virtualization features:
  Virtualization:         VT-x
  Hypervisor vendor:      Microsoft
  Virtualization type:    full
Caches (sum of all):
  L1d:                    192 KiB (6 instances)
  L1i:                    192 KiB (6 instances)
  L2:                     1.5 MiB (6 instances)
  L3:                     12 MiB (1 instance)
Vulnerabilities:
  Gather data sampling:   Unknown: Dependent on hypervisor status
  Itlb multihit:          KVM: Mitigation: VMX disabled
  L1tf:                   Not affected
  Mds:                    Not affected
  Meltdown:               Not affected
  Mmio stale data:        Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown
  Reg file data sampling: Not affected
  Retbleed:               Mitigation; Enhanced IBRS
  Spec rstack overflow:   Not affected
  Spec store bypass:      Mitigation; Speculative Store Bypass disabled via prctl and seccomp
  Spectre v1:             Mitigation; usercopy/swapgs barriers and __user pointer sanitization
  Spectre v2:             Mitigation; Enhanced / Automatic IBRS; IBPB conditional; RSB filling; PBRSB-eIBRS SW sequence;
                           BHI SW loop, KVM SW loop
  Srbds:                  Unknown: Dependent on hypervisor status
  Tsx async abort:        Not affected

解决建议

1. 移除同步IO操作

线程中频繁调用的cout是全局同步的IO操作,会导致线程阻塞等待输出锁,完全抵消多线程并行优势。注释掉所有cout语句后再测试性能:

void* threadnfunc (void *arg) {
    int j = *(int*)arg;
    // cout << "Starting thread " << j << endl;
    for(int i=0; i<N; i++) x[j]+=1;
    // cout << "Thread " << j << " execution completed!" << endl;
    return NULL;
}

2. 调整WSL2配置与主机电源计划

  • 将Windows主机电源计划设置为高性能,避免系统限制CPU性能
  • 关闭WSL后修改%USERPROFILE%\.wslconfig,确保配置合理:
    [wsl2]
    processors=12
    memory=8GB  # 根据物理内存分配,避免内存不足触发交换
    nestedVirtualization=false
    
    执行wsl --shutdown后重启WSL生效

3. 匹配线程数与物理核心数

你的CPU有6个物理核心,超线程带来的12个逻辑核心共享物理资源。将num_threads改为6(等于物理核心数),减少逻辑核心间的资源竞争,测试性能是否提升。

4. 强制CPU高性能模式

在WSL2中安装cpupower工具并强制CPU运行在最高频率:

sudo apt install linux-tools-common linux-tools-generic
sudo cpupower frequency-set --governor performance

5. 避免伪共享问题

全局数组x的元素可能共享CPU缓存行,导致缓存无效化开销。用缓存行对齐修饰数组:

long x[num_threads] __attribute__((aligned(64))) = {0};

内容的提问来源于stack exchange,提问作者Omair Siddique

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 01:37:35