You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何绑定繁忙CPU核心会导致OpenMP程序性能骤降?

问题背景

我的Linux系统配备12个CPU核心,其中核心2-11已被隔离,核心0和1被其他程序占用至接近100%,其余核心处于空闲状态。

第一轮测试

执行命令:

export GOMP_CPU_AFFINITY=2,3,4
export OMP_NUM_THREADS=3
taskset -c $GOMP_CPU_AFFINITY perf stat -d ./test_openmp

测试输出:

Performance counter stats for './test_openmp':

         47,654.74 msec task-clock:u              #    2.981 CPUs utilized           
                 0      context-switches:u        #    0.000 /sec                   
                 0      cpu-migrations:u          #    0.000 /sec                   
           115,358      page-faults:u             #    2.421 K/sec                  
   159,245,881,934      cycles:u                  #    3.342 GHz                    
   250,009,309,156      instructions:u            #    1.57  insn per cycle         
    20,002,132,172      branches:u                #  419.730 M/sec                  
           117,268      branch-misses:u           #    0.00% of all branches        
   110,002,614,320      L1-dcache-loads:u         #    2.308 G/sec                  
    10,796,435,741      L1-dcache-load-misses:u   #    9.81% of all L1-dcache accesses
                 0      LLC-loads:u               #    0.000 /sec                   
                 0      LLC-load-misses:u         #    0.00% of all LL-cache accesses

      15.986638336 seconds time elapsed

      47.175831000 seconds user
       0.414928000 seconds sys

第二轮测试

执行命令:

export GOMP_CPU_AFFINITY=1,2,3,4
export OMG_NUM_THREADS=4

taskset -c $GOMP_CPU_AFFINITY perf stat -d ./test_openmp

注:此处OMG_NUM_THREADS应为笔误,实际需设置OMP_NUM_THREADS=4才能正确指定OpenMP线程数

测试输出:

pid: 4118342

 Performance counter stats for './test_openmp':

         48,241.03 msec task-clock:u              #    1.072 CPUs utilized           
                 0      context-switches:u        #    0.000 /sec                   
                 0      cpu-migrations:u          #    0.000 /sec                   
           119,879      page-faults:u             #    2.485 K/sec                  
   161,605,704,451      cycles:u                  #    3.350 GHz                    
   250,011,376,400      instructions:u            #    1.55  insn per cycle         
    20,002,726,448      branches:u                #  414.641 M/sec                  
           118,657      branch-misses:u           #    0.00% of all branches        
   110,002,938,510      L1-dcache-loads:u         #    2.280 G/sec                  
    10,796,444,713      L1-dcache-load-misses:u   #    9.81% of all L1-dcache accesses
                 0      LLC-loads:u               #    0.000 /sec                   
                 0      LLC-load-misses:u         #    0.00% of all LL-cache accesses

      45.012033357 seconds time elapsed

      47.764469000 seconds user
       0.399934000 seconds sys

技术疑问

为何增加绑定一个繁忙的核心1后,OpenMP程序的运行时间大幅延长,CPU利用率反而显著降低?

测试代码

#include <iostream>
#include <cstdint>
#include <unistd.h>

constexpr int64_t N = 100000;
int m = N;
int n = N;

int main() {
  double* a = new double[N];
  double* c = new double[N];
  double* b = new double[N*N];

  std::cout << "pid: " << getpid() << std::endl;

#pragma omp parallel for default(none) shared(m,n,a,b,c)

for (int i=0; i<m; i++) {
 double sum = 0.0;
 for (int j=0; j<n; j++)
   sum += b[i+j*N]*c[j];
   a[i] = sum;
}

  return 0;
}

问题分析

核心原因是绑定到核心1的OpenMP线程几乎得不到CPU时间片,具体拆解:

  1. 线程调度被抢占
    核心1已被其他程序占满(CPU使用率接近100%),操作系统调度器几乎不会给绑定到该核心的OpenMP线程分配运行时间。虽然设置了4个线程,但实际只有3个线程(绑定到2、3、4核心)在正常工作,第4个线程完全处于等待调度的状态。

  2. CPU利用率的计算逻辑
    perf显示的CPUs utilized是用线程累计运行时间(task-clock)除以程序实际耗时(time elapsed):

  • 第一轮测试:3个线程跑满,累计task-clock约47.6秒,实际耗时约16秒,利用率≈47.6/16≈2.98,接近3核心满负载。
  • 第二轮测试:累计task-clock仅增加到约48.2秒(说明实际完成的计算量和第一轮几乎一致),但实际耗时拉长到约45秒,利用率≈48.2/45≈1.07——总耗时被拖长后,有效工作的3个线程分摊到更长时间里,利用率自然下降。
  1. 程序总耗时被拖慢
    OpenMP的parallel for需要等待所有线程完成任务才会结束。绑定到核心1的线程进度极慢,整个程序必须等它"完成",导致总耗时从16秒暴涨到45秒。

  2. 关键验证数据
    两轮测试的user时间几乎一致(第一轮47.17秒,第二轮47.76秒),说明程序实际完成的计算量没有变化,只是第二轮把相同的计算量分摊到了更长的实际时间中,完全是核心1的线程拖了后腿。

内容的提问来源于stack exchange,提问作者Xiaoyong Guo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 09:39:53