WSL2中C++多线程程序无性能提升的问题排查与解决
WSL2中多线程pthread程序无性能提升的解决办法
问题场景
我在WSL2中运行以下基于pthread的C++多线程程序:
#include <iostream> #include <pthread.h> #define N 1024*1024*1024 #define num_threads 8 using namespace std; long x[num_threads] = {0}; // thread function prototype void* threadnfunc (void *); void func (int); int main() { cout << "With threads" << endl; pthread_t my_threads[num_threads]; int arg_to_be_passed_to_func[num_threads]; // idx of array x // create threads for(int i=0; i<num_threads; i++){ arg_to_be_passed_to_func[i] = i; pthread_create (&my_threads[i], NULL, threadnfunc, &arg_to_be_passed_to_func[i]); } // start threads for(int i=0; i<num_threads; i++){ pthread_join (my_threads[i], NULL); } // cout << "Without threads" << endl; // for(int i=0; i<num_threads; i++){ // func(i); // } return 0; } void* threadnfunc (void *arg) { int j = *(int*)arg; cout << "Starting thread " << j << endl; for(int i=0; i<N; i++) x[j]+=1; cout << "Thread " << j << " execution completed!" << endl; return NULL; } void func(int j) { cout << "Starting thread " << j << endl; for(int i=0; i<N; i++) x[j]+=1; cout << "Thread " << j << " execution completed!" << endl; }
原本预期多线程版本相比单线程会显著提速,但在WSL2中两者执行时间几乎一致,而原生Ubuntu虚拟机上能看到明显加速效果。
系统信息
- 处理器:Intel Core i7-10750H
- 基准时钟频率:2.60GHz
- 内存:16GB
- 物理核心数:6
- 逻辑处理器数:12
WSL配置与CPU状态
- WSL内执行
nproc返回12,已通过%USERPROFILE%\.wslconfig指定processors=12,确认WSL可访问所有逻辑处理器 - BIOS中已开启虚拟化与超线程
- Windows任务管理器显示所有CPU(0-11)利用率为30-35%
cat /proc/cpuinfo | grep MHz显示每个CPU频率为2592MHz,接近基准时钟lscpu输出如下:
Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 39 bits physical, 48 bits virtual Byte Order: Little Endian CPU(s): 12 On-line CPU(s) list: 0-11 Vendor ID: GenuineIntel Model name: Intel(R) Core(TM) i7-10750H CPU @ 2.60GHz CPU family: 6 Model: 165 Thread(s) per core: 2 Core(s) per socket: 6 Socket(s): 1 Stepping: 2 BogoMIPS: 5184.01 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse s se2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon rep_good nopl xtopology cpuid pni pclmulqdq vmx ssse3 fma cx16 pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f1 6c rdrand hypervisor lahf_lm abm 3dnowprefetch invpcid_single ssbd ibrs ibpb stibp ibrs_enhanc ed tpr_shadow vnmi ept vpid ept_ad fsgsbase bmi1 avx2 smep bmi2 erms invpcid rdseed adx smap c lflushopt xsaveopt xsavec xgetbv1 xsaves md_clear flush_l1d arch_capabilities Virtualization features: Virtualization: VT-x Hypervisor vendor: Microsoft Virtualization type: full Caches (sum of all): L1d: 192 KiB (6 instances) L1i: 192 KiB (6 instances) L2: 1.5 MiB (6 instances) L3: 12 MiB (1 instance) Vulnerabilities: Gather data sampling: Unknown: Dependent on hypervisor status Itlb multihit: KVM: Mitigation: VMX disabled L1tf: Not affected Mds: Not affected Meltdown: Not affected Mmio stale data: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown Reg file data sampling: Not affected Retbleed: Mitigation; Enhanced IBRS Spec rstack overflow: Not affected Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl and seccomp Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; RSB filling; PBRSB-eIBRS SW sequence; BHI SW loop, KVM SW loop Srbds: Unknown: Dependent on hypervisor status Tsx async abort: Not affected
解决建议
1. 移除同步IO操作
线程中频繁调用的cout是全局同步的IO操作,会导致线程阻塞等待输出锁,完全抵消多线程并行优势。注释掉所有cout语句后再测试性能:
void* threadnfunc (void *arg) { int j = *(int*)arg; // cout << "Starting thread " << j << endl; for(int i=0; i<N; i++) x[j]+=1; // cout << "Thread " << j << " execution completed!" << endl; return NULL; }
2. 调整WSL2配置与主机电源计划
- 将Windows主机电源计划设置为高性能,避免系统限制CPU性能
- 关闭WSL后修改
%USERPROFILE%\.wslconfig,确保配置合理:
执行[wsl2] processors=12 memory=8GB # 根据物理内存分配,避免内存不足触发交换 nestedVirtualization=falsewsl --shutdown后重启WSL生效
3. 匹配线程数与物理核心数
你的CPU有6个物理核心,超线程带来的12个逻辑核心共享物理资源。将num_threads改为6(等于物理核心数),减少逻辑核心间的资源竞争,测试性能是否提升。
4. 强制CPU高性能模式
在WSL2中安装cpupower工具并强制CPU运行在最高频率:
sudo apt install linux-tools-common linux-tools-generic sudo cpupower frequency-set --governor performance
5. 避免伪共享问题
全局数组x的元素可能共享CPU缓存行,导致缓存无效化开销。用缓存行对齐修饰数组:
long x[num_threads] __attribute__((aligned(64))) = {0};
内容的提问来源于stack exchange,提问作者Omair Siddique
相关产品推荐
相关产品推荐

