You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

线程池在低CPU核数机器中延迟过高的原因排查

线程池任务延迟差异问题分析

我自行实现了一个线程池并行执行任务,测试时发现了一个奇怪的现象:在16核i9-10900K物理机上,每个任务的延迟仅约30us;但在2核AWS虚拟机上,延迟超过10000us。我清楚2核搭配16线程无法真正并行,但延迟涨幅远超预期,想请教该问题中的耗时环节在哪里?

线程池实现代码(thread_pool.hpp)

#pragma once
#include <vector>
#include <queue>
#include <memory>
#include <thread>
#include <mutex>
#include <condition_variable>
#include <future>
#include <functional>
#include <stdexcept>

class ThreadPool {
public:
  ThreadPool(size_t threads) : stop(false) {
    for(size_t i = 0; i < threads; ++i)
      workers.emplace_back(  // thread variable, emplace back just pass in a function for thread to construct
        [this] {
          for(;;) {
            std::function<void()> task;
            {   
              std::unique_lock<std::mutex> lock(this->queue_mutex);
              this->condition.wait(lock, [this] { return this->stop || !this->tasks.empty(); }); 
              if(this->stop && this->tasks.empty()) return;  // only on stop and empty task, perfect exit
              task = std::move(this->tasks.front());  // get the first task
              this->tasks.pop();
            }   
            task();
          }   
        }); 
  }


  template<class F, class... Args>
  auto enqueue(F&& f, Args&&... args) -> std::future<typename std::result_of<F(Args...)>::type> {
    using return_type = typename std::result_of<F(Args...)>::type;
    auto task = std::make_shared< std::packaged_task<return_type()> >(std::bind(std::forward<F>(f), std::forward<Args>(args)...));
    std::future<return_type> res = task->get_future();
    {   
      std::unique_lock<std::mutex> lock(queue_mutex);
      // don't allow enqueueing after stopping the pool
      if(stop) throw std::runtime_error("enqueue on stopped ThreadPool");
      tasks.emplace([task](){ (*task)(); }); 
    }   
    condition.notify_one();
    return res;
  }

  ~ThreadPool() {
    {   
      std::unique_lock<std::mutex> lock(queue_mutex);
      stop = true;
    }   
    condition.notify_all();
    for(std::thread &worker: workers) worker.join();
  }
private:
    std::vector< std::thread > workers;
    std::queue< std::function<void()> > tasks;
    std::mutex queue_mutex;
    std::condition_variable condition;
    bool stop;
};

测试代码(main.cpp)

#include "./thread_pool.hpp"
#include <sys/time.h>

int main() {
  ThreadPool tp(16);
  timeval t;
  for (size_t i = 0; i < 16; ++i) {
    gettimeofday(&t, NULL);
    tp.enqueue([t, i]() {
      timeval t1; 
      gettimeofday(&t1, NULL);
      int lat = (t1.tv_sec - t.tv_sec) * 1000000 + t1.tv_usec - t.tv_usec;
      printf("this is loop %zu, lat = %d\n", i, lat);
    }); 
  }
}

核心耗时环节分析

  • 线程上下文切换风暴:2核机器运行16个线程,每个线程都需要抢占CPU时间片,频繁的上下文切换会带来巨大开销——每次切换需要保存/恢复寄存器状态、刷新TLB(翻译后备缓冲器),这些操作单次就会消耗数微秒到数十微秒,16个线程反复切换的累积开销,直接拉高了每个任务的等待时间。

  • 锁竞争与条件变量唤醒低效:线程池使用单一把queue_mutex保护任务队列,在2核+16线程的场景下,16个工作线程会频繁竞争这把锁来获取任务,导致大量线程阻塞在锁等待上。同时,条件变量的唤醒机制在CPU资源紧张时,被唤醒的线程无法立即获得CPU执行权,需要等待调度,进一步拉长了从任务提交到执行的延迟。

  • 云虚拟机的额外调度开销:AWS虚拟机本身依赖物理CPU的分时调度,相比本地物理机,虚拟机之间的CPU竞争会带来额外的调度延迟。当2核虚拟机内的16个线程争夺CPU时,还要和同一物理节点上的其他虚拟机竞争资源,这会让任务的等待时间进一步放大。

  • 任务排队延迟:16个任务要在2核上执行,每个核心需要处理8个任务,任务必须排队等待CPU调度执行,这部分排队时间是延迟暴涨的核心组成部分之一。

内容的提问来源于stack exchange,提问作者kevin h

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 12:05:55