You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何C++互斥锁会严重影响多线程效率?附测试分析

多线程累加场景下互斥锁的性能瓶颈问题

我编写了一段用于测试多线程性能的代码,在循环中执行耗时计算后累加结果并统计执行时间。仅在累加结果的一行添加了互斥锁,但这一行锁直接拖垮了多线程性能,想知道原因。代码使用g++ -O3选项编译,同时测量了互斥锁的加锁/解锁耗时。

#include <chrono>
#include <cmath>
#include <functional>
#include <iomanip>
#include <iostream>
#include <mutex>
#include <vector>
#include <thread>

long double store;
std::mutex lock;

using ftype=std::function<long double(long int)>;
using loop_type=std::function<void(long int, long int, ftype)>;


///simple class to time the execution and print result.
struct time_n_print
{
  time_n_print() : 
    start(std::chrono::high_resolution_clock::now())
  {}
  
  ~time_n_print()
  {
    auto elapsed = std::chrono::high_resolution_clock::now() - start;
    auto ms = std::chrono::duration_cast<std::chrono::microseconds>(elapsed);
    std::cout << "Elapsed(ms)=" << std::setw(7) << ms.count();
    std::cout << "; Result: " << (long int)(store);
  }
  std::chrono::high_resolution_clock::time_point start;
};//class time_n_print

///do long and pointless calculations which result in 1.0
long double slow(long int i)
{
    long double pi=3.1415926536;
    long double i_rad  = (long double)(i) * pi / 180;
    long double sin_i  = std::sin(i_rad);
    long double cos_i  = std::cos(i_rad);
    long double sin_sq = sin_i * sin_i;
    long double cos_sq = cos_i * cos_i;
    long double log_sin_sq = std::log(sin_sq);
    long double log_cos_sq = std::log(cos_sq);
    sin_sq = std::exp(log_sin_sq);
    cos_sq = std::exp(log_cos_sq);
    long double sum_sq = sin_sq + cos_sq;
    long double result = std::sqrt(sum_sq);
    return result;
}

///just return 1
long double fast(long int)
{
    return 1.0;
}

///sum everything up with mutex
void loop_guarded(long int a, long int b, ftype increment)
{
  for(long int i = a; i < b; ++i)
  {
    long double inc = increment(i);
    {
      std::lock_guard<std::mutex> guard(lock);
      store += inc;
    }
  }
}//loop_guarded

///sum everything up without locks
void loop_unguarded(long int a, long int b, ftype increment)
{
  for(long int i = a; i < b; ++i)
  {
    long double inc = increment(i);
    {
      store += inc;
    }
  }
}//loop_unguarded

//run calculations on multiple threads.
void run_calculations(int size, 
                      int nthreads, 
                loop_type loop, 
                    ftype increment)
{
  store = 0.0;
  std::vector<std::thread> tv;
  long a(0), b(0);
  for(int n = 0; n < nthreads; ++n)
  {
    a = b;
    b = n < nthreads - 1 ? a + size / nthreads : size;
    tv.push_back(std::thread(loop, a, b, increment));
  }
  //Wait, until all threads finish
  for(auto& t : tv)
  {
    t.join();
  }
}//run_calculations

int main()
{
  long int size = 10000000;
  {
    std::cout << "\n1 thread  - fast, unguarded : ";
    time_n_print t;
    run_calculations(size, 1, loop_unguarded, fast);
  }
  {
    std::cout << "\n1 thread  - fast, guarded   : ";
    time_n_print t;
    run_calculations(size, 1, loop_guarded, fast);
  }
  std::cout << std::endl;
  {
    std::cout << "\n1 thread  - slow, unguarded : ";
    time_n_print t;
    run_calculations(size, 1, loop_unguarded, slow);
  }
  {
    std::cout << "\n2 threads - slow, unguarded : ";
    time_n_print t;
    run_calculations(size, 2, loop_unguarded, slow);
  }
  {
    std::cout << "\n3 threads - slow, unguarded : ";
    time_n_print t;
    run_calculations(size, 3, loop_unguarded, slow);
  }
  {
    std::cout << "\n4 threads - slow, unguarded : ";
    time_n_print t;
    run_calculations(size, 4, loop_unguarded, slow);
  }
  std::cout << std::endl;
  {
    std::cout << "\n1 thread  - slow, guarded   : ";
    time_n_print t;
    run_calculations(size, 1, loop_guarded, slow);
  }
  {
    std::cout << "\n2 threads - slow, guarded   : ";
    time_n_print t;
    run_calculations(size, 2, loop_guarded, slow);
  }
  {
    std::cout << "\n3 threads - slow, guarded   : ";
    time_n_print t;
    run_calculations(size, 3, loop_guarded, slow);
  }
  {
    std::cout << "\n4 threads - slow, guarded   : ";
    time_n_print t;
    run_calculations(size, 4, loop_guarded, slow);
  }
  std::cout << std::endl;
  return 0;
}

典型输出(4核Linux机器)

1 thread  - fast, unguarded : Elapsed(ms)=  32826; Result: 10000000  
1 thread  - fast, guarded   : Elapsed(ms)= 172208; Result: 10000000

1 thread  - slow, unguarded : Elapsed(ms)=2131659; Result: 10000000  
2 threads - slow, unguarded : Elapsed(ms)=1079671; Result: 9079646  
3 threads - slow, unguarded : Elapsed(ms)= 739284; Result: 8059758  
4 threads - slow, unguarded : Elapsed(ms)= 564641; Result: 7137484  

1 thread  - slow, guarded   : Elapsed(ms)=2198650; Result: 10000000  
2 threads - slow, guarded   : Elapsed(ms)=1468137; Result: 10000000  
3 threads - slow, guarded   : Elapsed(ms)=1306659; Result: 10000000  
4 threads - slow, guarded   : Elapsed(ms)=1549214; Result: 10000000

观察到的现象

  • 互斥锁的加锁/解锁耗时远高于long double的递增操作;
  • 无锁多线程的性能提升符合预期,但竞争条件导致累加结果大量丢失;
  • 加锁后,线程数超过2个时无性能提升,甚至4线程性能比2线程下降。

核心问题

为什么仅占总执行时间不到10%的锁代码段,会严重拖垮多线程性能?我知道可以通过线程局部累加最后汇总的方式解决,但想知道问题的根源。

更新:感谢各位回答与评论。本质原因是:如果每个线程有7-8%的时间处于锁定状态,就无法获得良好的性能提升。若在slow函数中添加10次循环,带锁与无锁版本的4线程性能提升完全一致。我的经验法则是:锁定状态的耗时占比不应超过总执行时间的1%。


问题根源解析

  1. 锁的开销不止加解锁指令:互斥锁的加解锁操作,在锁被占用时会触发线程从用户态切换到内核态等待,这个上下文切换的开销远大于锁本身的指令耗时。即使单次锁操作占比不高,但高频的锁竞争会让大量线程频繁切换状态,整体开销被急剧放大。
  2. 并行逻辑被锁串行化:代码中每次计算后都要加锁更新全局变量,这相当于多线程在锁操作这一步被迫串行执行。当线程数超过2时,新增线程的计算能力完全被锁等待的开销抵消,甚至因为更多上下文切换导致性能下降。
  3. 阿姆达尔定律的限制:根据阿姆达尔定律,程序的最大加速比由串行部分的占比决定。假设锁相关的串行部分占比8%,理论上4线程的最大加速比为1/(0.08 + (1-0.08)/4) ≈ 3.22,但实际中因锁竞争的额外开销,这个理论值还无法达到,甚至出现加速比不升反降的情况。当你放大slow函数的计算量,锁的占比降到0.8%左右时,串行部分的影响可以忽略,多线程就能发挥正常的并行能力。

内容的提问来源于stack exchange,提问作者one_two_three

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 07:25:24