std::async执行的任务能否复用空闲线程?技术问询
你观察到的现象其实和std::async的默认行为以及Windows平台下的标准库实现直接相关,我来一步步拆解清楚:
首先得明确std::async的默认启动策略:当你不指定std::launch参数时,它采用的是std::launch::async | std::launch::deferred的组合策略——这意味着标准库有权选择两种执行方式之一:要么立刻创建新线程跑任务,要么延迟到你调用future.get()或future.wait()时才执行(甚至可能直接在调用线程上同步跑)。但在Windows的MSVC标准库实现里,对于你这种只创建future却不立刻等待的场景,默认会走std::launch::async路线,也就是每个任务都新建一个独立线程。
为什么会出现161个等待线程?
你的代码开了160个任务,每个任务对应一个线程,再加上你的主线程,刚好就是161个线程。这些线程因为任务逻辑是std::this_thread::sleep_for(20秒),所以全部进入休眠等待状态,这就是你看到的结果。
这个线程数量是不是过多?
对于这种纯休眠、没有实际计算逻辑的任务来说,确实非常浪费。每个线程都要占用内核对象、默认1MB左右的栈空间等系统资源,160个线程会消耗不少内存和内核资源,完全是没必要的开销。
为什么空闲线程没法被复用?
核心原因是:C++标准并没有要求std::async的async策略必须使用线程池复用线程,不同标准库的实现差异很大:
- 像MSVC的标准库,在默认策略下每个
std::async任务都会新建线程,不会复用已有线程; - 而GCC的libstdc++在某些版本里会用线程池来复用线程,但这属于实现细节,不是标准强制要求的。
想要线程复用该怎么做?
如果希望复用线程来处理大量任务,最靠谱的方式是显式使用线程池,有几种选择:
- 自己封装一个简单的线程池(比如下面的示例框架);
- 使用C++20的
std::jthread配合任务队列实现; - 借助第三方线程池库(比如Boost.ThreadPool)。
这里给你一个极简的线程池实现框架参考:
#include <vector> #include <queue> #include <thread> #include <mutex> #include <condition_variable> #include <functional> class ThreadPool { public: // 用硬件核心数初始化线程池大小,避免创建过多线程 ThreadPool() : ThreadPool(std::thread::hardware_concurrency()) {} explicit ThreadPool(size_t threadCount) { for (size_t i = 0; i < threadCount; ++i) { workers.emplace_back([this] { while (true) { std::function<void()> task; { std::unique_lock<std::mutex> lock(queueMutex); // 等待任务或停止信号 condition.wait(lock, [this] { return stop || !tasks.empty(); }); if (stop && tasks.empty()) return; // 取出任务 task = std::move(tasks.front()); tasks.pop(); } // 执行任务 task(); } }); } } // 提交任务到线程池 template<class F> void enqueue(F&& func) { { std::unique_lock<std::mutex> lock(queueMutex); tasks.emplace(std::forward<F>(func)); } condition.notify_one(); // 唤醒一个空闲线程处理任务 } // 析构时等待所有线程完成 ~ThreadPool() { { std::unique_lock<std::mutex> lock(queueMutex); stop = true; } condition.notify_all(); // 唤醒所有线程 for (auto& worker : workers) { worker.join(); } } // 禁用拷贝和移动 ThreadPool(const ThreadPool&) = delete; ThreadPool& operator=(const ThreadPool&) = delete; ThreadPool(ThreadPool&&) = delete; ThreadPool& operator=(ThreadPool&&) = delete; private: std::vector<std::thread> workers; std::queue<std::function<void()>> tasks; std::mutex queueMutex; std::condition_variable condition; bool stop = false; }; // 使用示例 int main() { ThreadPool pool; // 默认用硬件核心数创建线程 auto idleTask = []() { std::this_thread::sleep_for(std::chrono::seconds(20)); }; const int taskCount = 160; for (int i = 0; i < taskCount; ++i) { pool.enqueue(idleTask); } // 这里可以添加逻辑等待所有任务完成,比如用std::future追踪每个任务 return 0; }
用这种方式,线程数量会被限制在硬件核心数(比如8核就是8个线程),所有160个任务都会复用这些线程,不会出现大量闲置线程浪费资源的情况。
内容的提问来源于stack exchange,提问作者Avega

