线程池在低CPU核数机器中延迟过高的原因排查
线程池任务延迟差异问题分析
我自行实现了一个线程池并行执行任务,测试时发现了一个奇怪的现象:在16核i9-10900K物理机上,每个任务的延迟仅约30us;但在2核AWS虚拟机上,延迟超过10000us。我清楚2核搭配16线程无法真正并行,但延迟涨幅远超预期,想请教该问题中的耗时环节在哪里?
线程池实现代码(thread_pool.hpp)
#pragma once #include <vector> #include <queue> #include <memory> #include <thread> #include <mutex> #include <condition_variable> #include <future> #include <functional> #include <stdexcept> class ThreadPool { public: ThreadPool(size_t threads) : stop(false) { for(size_t i = 0; i < threads; ++i) workers.emplace_back( // thread variable, emplace back just pass in a function for thread to construct [this] { for(;;) { std::function<void()> task; { std::unique_lock<std::mutex> lock(this->queue_mutex); this->condition.wait(lock, [this] { return this->stop || !this->tasks.empty(); }); if(this->stop && this->tasks.empty()) return; // only on stop and empty task, perfect exit task = std::move(this->tasks.front()); // get the first task this->tasks.pop(); } task(); } }); } template<class F, class... Args> auto enqueue(F&& f, Args&&... args) -> std::future<typename std::result_of<F(Args...)>::type> { using return_type = typename std::result_of<F(Args...)>::type; auto task = std::make_shared< std::packaged_task<return_type()> >(std::bind(std::forward<F>(f), std::forward<Args>(args)...)); std::future<return_type> res = task->get_future(); { std::unique_lock<std::mutex> lock(queue_mutex); // don't allow enqueueing after stopping the pool if(stop) throw std::runtime_error("enqueue on stopped ThreadPool"); tasks.emplace([task](){ (*task)(); }); } condition.notify_one(); return res; } ~ThreadPool() { { std::unique_lock<std::mutex> lock(queue_mutex); stop = true; } condition.notify_all(); for(std::thread &worker: workers) worker.join(); } private: std::vector< std::thread > workers; std::queue< std::function<void()> > tasks; std::mutex queue_mutex; std::condition_variable condition; bool stop; };
测试代码(main.cpp)
#include "./thread_pool.hpp" #include <sys/time.h> int main() { ThreadPool tp(16); timeval t; for (size_t i = 0; i < 16; ++i) { gettimeofday(&t, NULL); tp.enqueue([t, i]() { timeval t1; gettimeofday(&t1, NULL); int lat = (t1.tv_sec - t.tv_sec) * 1000000 + t1.tv_usec - t.tv_usec; printf("this is loop %zu, lat = %d\n", i, lat); }); } }
核心耗时环节分析
线程上下文切换风暴:2核机器运行16个线程,每个线程都需要抢占CPU时间片,频繁的上下文切换会带来巨大开销——每次切换需要保存/恢复寄存器状态、刷新TLB(翻译后备缓冲器),这些操作单次就会消耗数微秒到数十微秒,16个线程反复切换的累积开销,直接拉高了每个任务的等待时间。
锁竞争与条件变量唤醒低效:线程池使用单一把
queue_mutex保护任务队列,在2核+16线程的场景下,16个工作线程会频繁竞争这把锁来获取任务,导致大量线程阻塞在锁等待上。同时,条件变量的唤醒机制在CPU资源紧张时,被唤醒的线程无法立即获得CPU执行权,需要等待调度,进一步拉长了从任务提交到执行的延迟。云虚拟机的额外调度开销:AWS虚拟机本身依赖物理CPU的分时调度,相比本地物理机,虚拟机之间的CPU竞争会带来额外的调度延迟。当2核虚拟机内的16个线程争夺CPU时,还要和同一物理节点上的其他虚拟机竞争资源,这会让任务的等待时间进一步放大。
任务排队延迟:16个任务要在2核上执行,每个核心需要处理8个任务,任务必须排队等待CPU调度执行,这部分排队时间是延迟暴涨的核心组成部分之一。
内容的提问来源于stack exchange,提问作者kevin h
相关产品推荐
相关产品推荐

